![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) Wenxuan Tanandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
0f587e80d3
|
Use Tensor Core Decode when gqa group size >= 4 (#8624)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-08-22 23:25:15 +08:00 |
|
 Wenxuan TanandLiangsheng Yin
|
0305c5053f
|
Reduce memory accumulation in long-running server (#8306)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2025-08-03 15:03:16 +08:00 |
|
Wenxuan Tan
|
a968c888c0
|
Fix torchvision version for Blackwell (#7015)
|
2025-06-09 15:50:19 -07:00 |
|
Wenxuan Tan
|
c429919def
|
misc: cache is_hopper_arch (#6799)
|
2025-06-01 15:28:31 -07:00 |
|
Wenxuan Tan
|
844a8f42c7
|
Fix LoRA bench (#6719)
|
2025-05-28 16:38:55 -07:00 |
|
Wenxuan Tan
|
66324895c6
|
[docs] Fix torch version (#6472)
|
2025-05-20 10:53:14 -07:00 |
|
Wenxuan Tan
|
22da3d978f
|
Fix "Avoid computing lse in Ragged Prefill when there's no prefix match" (#5555)
|
2025-05-05 10:32:17 -07:00 |
|
Wenxuan Tan
|
dfb322642f
|
Use device_id in dist init to reduce NCCL communicator warmup & creation overhead (#5728)
|
2025-04-26 18:11:09 -07:00 |
|
 Wenxuan TanandBaizhou Zhang
|
bfa3922451
|
Avoid computing lse in Ragged Prefill when there's no prefix. (#5476)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-04-18 01:13:57 -07:00 |
|
Wenxuan Tan
|
718c391fd7
|
[Hoxfix] Fix incomplete token_to_kv_pool refactor (#4121)
|
2025-03-05 19:32:42 -08:00 |
|
 Wenxuan Tanandyinfan98
|
0af1d239cb
|
[Docs] Add quantization docs (#3410)
Co-authored-by: yinfan98 <1106310035@qq.com>
|
2025-02-10 02:16:21 +08:00 |
|