Commit Graph
7199 Commits
Author SHA1 Message Date
fzyzcjy 64a11303ce Fix update weight error for blackwell DeepGEMM (#13910) 2025-11-25 13:28:12 -08:00
Lianmin Zheng 1ab6ce0e62 [Auto Sync] Improve profilers and simplify bench_one_batch_server.py (#13866) 2025-11-25 12:13:31 -08:00
Kangyan-Zhou e99ca6ac74 Improve nightly tests (#13903) 2025-11-25 12:04:49 -08:00
Jan Bernlöhr fcccaf9001 Add Llama4 attention backend auto-selection (#13421)
Signed-off-by: jbernloehr <jbernloehr@nvidia.com>
2025-11-25 11:54:21 -08:00
Tova Movshovitz 215a97fa6c Fix docstrings for v1 HiCacheStorage methods (#13851) 2025-11-25 11:48:26 -08:00
YAMYandhlu1 5eed5fc0b0 [DeepSeekV3.2] Centralize NSA dispatch logic in NativeSparseAttnBackend (#13544)
Co-authored-by: hlu1 <14827759+hlu1@users.noreply.github.com>
2025-11-25 11:32:30 -08:00
Baizhou Zhang 808b6dfdea [Minor] Fix lint (#13938) 2025-11-25 10:57:23 -08:00
Simo Lin 4852aa054c [misc] add llama3.1 chat template (#13935) 2025-11-25 09:31:54 -08:00
Zaili Wang f922bfd520 [CPU] Apply PR gating rule in CI workflow (#13933) 2025-11-26 00:44:16 +08:00
Mick dfd7ab9682 [diffusion] feat: support LoRA (#13859) 2025-11-26 00:21:33 +08:00
Mick 46673b4224 [diffusion] doc: add doc for LoRA usage (#13931) 2025-11-26 00:02:14 +08:00
Liangsheng Yin 3421d049aa [CI] rename: per_commit -> registered (#13928) 2025-11-25 23:05:06 +08:00
Liangsheng Yinandalisonshao d3d404d3d7 [CI] CI registry update (#13927)
Co-authored-by: alisonshao <54658187+alisonshao@users.noreply.github.com>
2025-11-25 22:43:19 +08:00
luchangli 64225a8ae9 fix nixl prefill crash make decode health check failed (#13657)
Signed-off-by: liluchang <liluchang@kingsoft.com>
2025-11-25 21:40:11 +08:00
Chen1022 d64bf6c6ce Support piecewise cuda graph for Qwen3-next (#13081) 2025-11-25 21:01:27 +08:00
LiwansiandEven Zhou 432ecf841e [Ascend] qwen optimization (#12078)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
2025-11-25 19:44:24 +08:00
Ke Bao 0b3f002daf Update release-whl-kernel.yml (#13921) 2025-11-25 19:05:38 +08:00
Mick 6f094deff0 [diffusion] CI: minor refactor CI for less code duplication (#13905) 2025-11-25 18:44:11 +08:00
ant-yy 59464dbf15 [Fix]: Further fix the buffer len of future map (#13916)
Signed-off-by: vito.yy <vito.yy@antgroup.com>
2025-11-25 18:09:28 +08:00
Yi ZhangandMick 1f7fcc10d5 [diffusion] profile: fix profiling bugs (#13642)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-11-25 17:58:35 +08:00
alisonshao dbab5d50a3 Add test_dummy_grok_models.py to not_in_ci section (#13908) 2025-11-25 16:51:25 +08:00
Lzhang-hub 760c20b360 update flashinfer_cubin==0.5.3 (#13848) 2025-11-25 00:10:34 -08:00
Xiaoyu Zhangandgithub-actions[bot] <github-actions[bot]@users.noreply.github.com> 407cb3ce1e [CI tiny fix] Enhance robustness of vision chunked prefill test with ROUGE-L metric (#13793)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2025-11-25 15:41:14 +08:00
alisonshao 7cc43bd453 Move test_dummy_grok_models.py from manual to srt (temporary) (#13901) 2025-11-24 23:36:53 -08:00
Zaili WangandFan Yin cce2d748ef remove RoPE CPU fp32 tests (#13827)
Co-authored-by: Fan Yin <1106310035@qq.com>
2025-11-24 23:22:35 -08:00
DarkSharpness c1dd9a9599 [Fix] JIT kernel dependencies in other platforms (#13889) 2025-11-24 23:19:17 -08:00
Douglas Yang ed8786b0b9 Adding nightly tests for Kimi-K2-thinking, Qwen3, minimax-m2, GLM4.6 (#13890) 2025-11-24 22:47:46 -08:00
alisonshao f9fe06309f Fix trace publish paths in nightly-test-nvidia workflow (#13888) 2025-11-24 21:58:26 -08:00
gongwei-130 8ff3ef1fef fix: draft model revision misuse model revision (#11893) 2025-11-24 21:13:37 -08:00
Yibo Cai da182e4b83 [CI] fix lint error (#13891) 2025-11-24 20:51:33 -08:00
alisonshaoMickgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
83e7207763 [diffusion] CI: add validation and cleanup for corrupted safetensors in multimodal loader (#13870)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-11-25 12:02:43 +08:00
Yibo Cai a2c388ba11 [CI] fix multimodel-gen-test job (#13874) 2025-11-24 19:46:23 -08:00
Mick 9384fa2729 [diffusion] refactor: remove training-related code (#13860) 2025-11-25 11:38:50 +08:00
alisonshao 173e73fa1e Fix nightly test job to fail when any test fails (#13871) 2025-11-24 19:33:46 -08:00
Siyuan Chen a164259efd [Router bugfix] Fix router_manager selecting the wrong router when enable-igw. (#13572) 2025-11-24 18:51:41 -08:00
b0a26ba624 Add support for bf16 x bf16 cutlass fused MoE (#10275)
Co-authored-by: Sam Li <lsam@nvidia.com>
Co-authored-by: jackeyhua <jackeyhuasjtu@gmail.com>
2025-11-24 18:49:39 -08:00
Binyao Jiang de430b6745 [Performance] Replace preprocess_video logic from GLM multimodal processor with transformer impl for speed up (up to 27% faster) and addressing OOM (up to 50x improvements) (#13487) 2025-11-24 18:17:13 -08:00
Qiaolin Yu 4b45d556a7 Overlap glm moe gemms in two cuda streams (#13786) 2025-11-24 18:15:24 -08:00
Even Zhouandc30031083 db0ffc09ef [NPU] Fix NPU CI (#13834)
Co-authored-by: c30031083 <chenxu140@huawei.com>
2025-11-25 10:09:36 +08:00
Glen Liu eb1d885400 add LoRA warning if loading a preexisting LoRA adapter with a different name (#13822) 2025-11-24 15:16:41 -08:00
Cheng Wan bf10869203 [Doc] Add an Introduction to Expert Parallelism (#13783) 2025-11-24 14:46:51 -08:00
Lianmin Zhenggithub-actions[bot] <github-actions[bot]@users.noreply.github.com>Hanming LuHanming Lu
e83bd1fadc [Auto Sync] Update schedule_batch.py, schedule_policy.py, b... (20251122) (#13763)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: Hanming Lu <hanming@x.ai>
2025-11-24 14:33:31 -08:00
gongwei-130 9dc15d8569 fix xgrammar_backend crash with malformed inputs (#13752) 2025-11-24 13:54:19 -08:00
Minglei Zhu fafaa2ccea [BugFix] fix outplace_fused_experts missing is_gated (#13864) 2025-11-24 12:36:14 -08:00
Simo Lin 9b4b344115 [model-gateway] add grpc server code owner (#13865) 2025-11-24 12:29:57 -08:00
fzyzcjy 94216a9cc4 Fix quantized moe checker fail for Qwen3 dense fp8 model (#13853) 2025-11-24 11:16:50 -08:00
Yuhao YaoandFan Yin 9535015d05 [Perf] Optimize DeepSeek-R1 w4afp8 glue kernels (#10027)
Co-authored-by: Fan Yin <1106310035@qq.com>
2025-11-24 11:05:38 -08:00
a3b578fc60 [model-gateway] Refactor router e2e responses tests (#13745)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
2025-11-24 10:58:53 -08:00
Liangsheng Yin b60e769d0e Tiny unpin uvloop for other backends (#13858) 2025-11-25 01:34:57 +08:00
Zhi Yiliuandlzy a95a38078b [Fix] Fix uvloop get_event_loop() is not suitable for 0.22.x (#13612)
Signed-off-by: lzy <tomlzy213@gmail.com>
Co-authored-by: lzy <tomlzy213@gmail.com>
2025-11-25 01:20:00 +08:00