yinghui
|
c23eda8589
|
Fix incorrect KV indices creation when page_size=32 in TRTLLM MLA backend (#11985)
|
2025-10-22 22:44:45 -07:00 |
|
Qiaolin Yu
|
d9a20fd28a
|
Use trtllm_mla decode kernel for draft extend in speculative decoding (#11664)
|
2025-10-21 11:42:09 +08:00 |
|
Elfie Guo
|
bebd0576e5
|
Integrate trtllm ragged attention for prefill self-attention (#9801)
|
2025-09-05 17:18:00 +08:00 |
|
Faraz
|
ff9b561817
|
Fix TRTLLM MLA Cuda KV Blocks Causing accuracy drop (#9675)
|
2025-08-29 17:16:10 -07:00 |
|
Faraz
|
f508cd3cb7
|
TRTLLM-MLA FP8 path (#8638)
Signed-off-by: Faraz Khoubsirat <58580514+farazkh80@users.noreply.github.com>
|
2025-08-11 14:02:13 -07:00 |
|
Faraz
|
4b04998d38
|
TRTLLM Gen MLA Decode Kernel Integration (same as #7938) (#8632)
Signed-off-by: Faraz Khoubsirat <58580514+farazkh80@users.noreply.github.com>
|
2025-07-31 16:03:40 -07:00 |
|