Commit Graph

9043 Commits

Author SHA1 Message Date
Julian Huang
db2425a00b [Fix]: correctly fetch ds32 config in tuning_fused_moe_triton (#17409)
Co-authored-by: 墨楼 <huangzhilin.hzl@antgroup.com>
2026-01-20 20:08:28 +08:00
amote-i
603f386c6b update docs of Ascend plateform (#17358) 2026-01-20 19:31:23 +08:00
ympcMark
f7a5e425c3 [3/N] Achieve fault tolerance at the DP level (#11657)
Co-authored-by: UNIDY <unidy2002@outlook.com>
Co-authored-by: Hank Han <hanhan7630@outlook.com>
2026-01-20 18:47:08 +08:00
Douglas Yang
d50dcd9b61 fix: Release pr pypi fix (#17382) 2026-01-19 23:00:54 -08:00
Yuan Luo
e6b7c04947 [Kimi-Linear] Refactor kimi-linear gate calculation to avoid duplicated code (#17160)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-01-20 14:29:24 +08:00
Michael
a3addd6203 [AMD] Add DeepSeek-V3.2 and VLMs model in nightly tests (#17179)
Co-authored-by: michaelzhang-ai <michaelzhang-ai@users.noreply.github.com>
Co-authored-by: YC Tseng <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-01-19 20:31:56 -08:00
Thomas Wang
6988a0f570 Disable mla persistent kernel when not using fp8 kv_cache (#17327) 2026-01-19 20:29:54 -08:00
Hexq0210
1b192cf198 [NPU] Update NPU doc for model and features supported (#17385) 2026-01-20 12:28:06 +08:00
Ratish P
c560e14216 [diffusion] fix: enforce 16-pixel alignment for flux2 to prevent shape mismatch crash (#17302) 2026-01-20 12:03:33 +08:00
Shangming Cai
71cb9d0302 [PD] Fix PP dynamic chunking with DP attention (#17339)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-01-20 11:27:13 +08:00
shuwenn
8fb45523f3 feat: support bitsandbytes quantization algorithm (#15325) 2026-01-19 18:36:56 -08:00
Raghav Ravishankar
84aef37808 fix function calling for Trinity (#17364) 2026-01-19 18:33:17 -08:00
Alison Shao
17c04b109d Enable parallel stage execution for scheduled CI runs (#16880) 2026-01-19 18:31:46 -08:00
Alison Shao
55c4288b3e Fix runner utilization workflow to use 24h default (#17378) 2026-01-19 18:31:33 -08:00
zijiexia
9f8b79f16f [Docs] Fix formatting in Evaluating New Models with SGLang (#17376) 2026-01-19 18:22:30 -08:00
b8zhong
7dc3cbe7ca [Docker] Fix CUDA 13 installing wrong nvidia-nccl-cu13 due to nixl-cu13 not breaking system package (#17370) 2026-01-20 09:52:30 +08:00
zijiexia
79ddc34c1c [Docs] Add new model evaluation docs (#17043)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
2026-01-19 16:35:03 -08:00
Aurick Qiao
09a9d214f7 Pipe customized_info through CudaGraphRunner output (#17088) 2026-01-19 15:48:49 -08:00
Alison Shao
7e40d52635 Move test_autoround.py to stage-b-test-large-1-gpu suite (#17336) 2026-01-19 14:36:19 -08:00
Alison Shao
057b07fc50 Disable unit-test-backend-4-gpu-gb200 job (#17367) 2026-01-19 14:35:58 -08:00
Douglas Yang
91d8c52d6d fix: updating pypi workflow with new base version formatting in pyproject.toml (#17366) 2026-01-19 13:23:11 -08:00
YC Tseng
e9a44ea607 [AMD] fix perf ci errors (#17363) 2026-01-19 13:04:15 -08:00
Hudson Xing
c1282da236 fix(ci): apply MMMU retry logic to all affected test files (#17329) 2026-01-19 10:00:29 -08:00
shuwenn
71279e31f7 [CI] fix test_vlm_models.py (#17049) 2026-01-19 10:00:00 -08:00
YC Tseng
1a053a810c [AMD] CI - add partitions for stage-b-test-small-1-gpu-amd (#17345) 2026-01-19 08:16:07 -08:00
Bingxu Chen
2ea02f0642 [AMD CI] Migrate and Add More Testcases (#17116)
Co-authored-by: yctseng0211 <yctseng@amd.com>
2026-01-19 08:07:39 -08:00
zhangheng
20b0523eca [RadixTree][1/N Refactor]: Support unified match_prefix params (#17142)
Co-authored-by: yizhang2077 <1109276519@qq.com>
Co-authored-by: pansicheng <sicheng.pan.chn@gmail.com>
2026-01-19 22:39:40 +08:00
Ke Bao
ce8a6ac690 Evict swa kv cache during decoding (#17220) 2026-01-19 22:36:52 +08:00
ybyang
ebca5879a1 Fix v32 continue_final_message not work (#16567) 2026-01-19 21:45:39 +08:00
Xiaoyu Zhang
cc410a1088 [Diffusion] Apply qknorm to flux2 and apply lightx2v rms_norm_one_pass kernel(without residual) (#17305)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-19 21:25:33 +08:00
b8zhong
f374623fa9 [Refactor] Set fp4-gemm-backend=auto on SM100 and rename fp4-gemm-backend with flashinfer_ prefix (#17309) 2026-01-19 20:09:07 +08:00
Shu Wang
5c02217746 Inclusion of nvfp4 blockscale in EPLB Rebalance (#17158) 2026-01-19 17:45:27 +08:00
Alison Shao
8916b9d080 Migrate performance, accuracy, and quantization tests to CI registry (#17177)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2026-01-18 23:25:24 -08:00
Xiaoyu Zhang
a3d9a21882 Revert "[Perf] fuse q, k norm for Flux2Attention (#17241)" (#17332) 2026-01-19 15:24:11 +08:00
Alison Shao
fb88fb672e fix(ci): rate limit and permission errors in trace publishing (#17238) 2026-01-18 23:20:22 -08:00
Alison Shao
2d72e168fd [CI] Add partition to stage-b-test-large-1-gpu (11->12) (#17245) 2026-01-18 23:19:39 -08:00
Minglei Zhu
64946679b5 [Perf] fuse q, k norm for Flux2Attention (#17241)
Co-authored-by: Minglei Zhu <zminglei@linkedin.com>
2026-01-19 14:33:34 +08:00
Kartik Ramesh
5836324c55 KV Cache Events with Attention DP bug fix (#16030) (#16412) 2026-01-19 14:13:48 +08:00
yudian0504
9fe56cd0fb Fix kernel selection in biased_grouped_topk_gpu (#17325) 2026-01-19 14:06:11 +08:00
Gaoji Liu
858a4d659b support new qwen3_coder_detector (#16744)
Co-authored-by: liugaoji.lgj <liugaoji.lgj@alibaba-inc.com>
2026-01-18 21:16:27 -08:00
Lianmin Zheng
e619f53113 [Auto Sync] Update tokenizer_manager.py (20260119) (#17317)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-01-18 20:32:19 -08:00
Lianmin Zheng
fc4b932f4e Update code sync scripts (#17319) 2026-01-18 19:57:28 -08:00
Yongfei Xu
d2105d4abd [DeepSeek v3.2] Opt MTP decode cuda batch sizes and nsa implementation (#16961) 2026-01-19 11:54:11 +08:00
Lee Nau
84c8390514 Use dsv3 optimized routing fused_topk_deepseek instead of moe_fused_gate (#15347) 2026-01-19 11:50:16 +08:00
Baizhou Zhang
ea879c7739 [Minor] Correct sglang version when installing from source (#17315) 2026-01-18 19:36:16 -08:00
Shangming Cai
0227db8926 [PD] Optimize MHA models pp util calculation logic (#17306) 2026-01-19 11:23:15 +08:00
Glen Liu
ad1b4e4728 [Feature] overlap LoRA weight loading with compute (#15512) 2026-01-19 10:43:17 +08:00
Mick
51f147ada3 Update CODEOWNERS for multimodal_gen (#17308)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
2026-01-19 08:39:20 +08:00
Koushik Dutta
d3eafc7357 [GLM 4.7] Add RTX 6000 Pro aka sm120 (#17235)
Co-authored-by: root <root@ubuntu-nvidia.localdomain>
2026-01-18 13:33:19 -08:00
Jinyan Chen
e00b43442d [jit-kernel] Add CuTe DSL GDN Decode Kernel (#15631)
Co-authored-by: Jinyan Chen <jinyanc@nvidia.com>
2026-01-18 12:54:36 -08:00