Commit Graph
11 Commits
Author SHA1 Message Date
Wenxuan Tanandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> 0f587e80d3 Use Tensor Core Decode when gqa group size >= 4 (#8624)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-08-22 23:25:15 +08:00
Wenxuan TanandLiangsheng Yin 0305c5053f Reduce memory accumulation in long-running server (#8306)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2025-08-03 15:03:16 +08:00
Wenxuan Tan a968c888c0 Fix torchvision version for Blackwell (#7015) 2025-06-09 15:50:19 -07:00
Wenxuan Tan c429919def misc: cache is_hopper_arch (#6799) 2025-06-01 15:28:31 -07:00
Wenxuan Tan 844a8f42c7 Fix LoRA bench (#6719) 2025-05-28 16:38:55 -07:00
Wenxuan Tan 66324895c6 [docs] Fix torch version (#6472) 2025-05-20 10:53:14 -07:00
Wenxuan Tan 22da3d978f Fix "Avoid computing lse in Ragged Prefill when there's no prefix match" (#5555) 2025-05-05 10:32:17 -07:00
Wenxuan Tan dfb322642f Use device_id in dist init to reduce NCCL communicator warmup & creation overhead (#5728) 2025-04-26 18:11:09 -07:00
Wenxuan TanandBaizhou Zhang bfa3922451 Avoid computing lse in Ragged Prefill when there's no prefix. (#5476)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2025-04-18 01:13:57 -07:00
Wenxuan Tan 718c391fd7 [Hoxfix] Fix incomplete token_to_kv_pool refactor (#4121) 2025-03-05 19:32:42 -08:00
Wenxuan Tanandyinfan98 0af1d239cb [Docs] Add quantization docs (#3410)
Co-authored-by: yinfan98 <1106310035@qq.com>
2025-02-10 02:16:21 +08:00