Commit Graph

7207 Commits

Author SHA1 Message Date
Yuan Luo
ca5c8b16f6 [VLM] Support InternVL Vision Encoder Data Parallelism (#13925)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-11-26 11:43:05 +08:00
timmy-feng
f33e5d1ef1 fix: spec overlap predict shape does not match verify output shapes (#12786) 2025-11-26 11:01:06 +08:00
Kangyan-Zhou
13e5beeab4 Fix Deepseek v3.1 loading issue (#13954) 2025-11-25 18:28:52 -08:00
sglang-bot
c53e729d45 chore: bump sgl-kernel version to 0.3.18.post1 (#13951)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2025-11-25 18:14:28 -08:00
Kangyan-Zhou
03a26557b7 Fix nightly-test-nvidia.yml to have the correct trigger (#13950) 2025-11-25 17:34:34 -08:00
Baizhou Zhang
873382a910 [Tiny]Upgrade README for sgl-kernel (#13945) 2025-11-25 16:46:30 -08:00
sglang-bot
391a863b3f chore: bump sgl-kernel version to 0.3.18.post1 (#13942)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2025-11-25 15:31:53 -08:00
Fan Yin
36b1bcd242 [chore] update torch version to 2.9 (#12969)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2025-11-25 14:47:34 -08:00
fzyzcjy
64a11303ce Fix update weight error for blackwell DeepGEMM (#13910) 2025-11-25 13:28:12 -08:00
Lianmin Zheng
1ab6ce0e62 [Auto Sync] Improve profilers and simplify bench_one_batch_server.py (#13866) 2025-11-25 12:13:31 -08:00
Kangyan-Zhou
e99ca6ac74 Improve nightly tests (#13903) 2025-11-25 12:04:49 -08:00
Jan Bernlöhr
fcccaf9001 Add Llama4 attention backend auto-selection (#13421)
Signed-off-by: jbernloehr <jbernloehr@nvidia.com>
2025-11-25 11:54:21 -08:00
Tova Movshovitz
215a97fa6c Fix docstrings for v1 HiCacheStorage methods (#13851) 2025-11-25 11:48:26 -08:00
YAMY
5eed5fc0b0 [DeepSeekV3.2] Centralize NSA dispatch logic in NativeSparseAttnBackend (#13544)
Co-authored-by: hlu1 <14827759+hlu1@users.noreply.github.com>
2025-11-25 11:32:30 -08:00
Baizhou Zhang
808b6dfdea [Minor] Fix lint (#13938) 2025-11-25 10:57:23 -08:00
Simo Lin
4852aa054c [misc] add llama3.1 chat template (#13935) 2025-11-25 09:31:54 -08:00
Zaili Wang
f922bfd520 [CPU] Apply PR gating rule in CI workflow (#13933) 2025-11-26 00:44:16 +08:00
Mick
dfd7ab9682 [diffusion] feat: support LoRA (#13859) 2025-11-26 00:21:33 +08:00
Mick
46673b4224 [diffusion] doc: add doc for LoRA usage (#13931) 2025-11-26 00:02:14 +08:00
Liangsheng Yin
3421d049aa [CI] rename: per_commit -> registered (#13928) 2025-11-25 23:05:06 +08:00
Liangsheng Yin
d3d404d3d7 [CI] CI registry update (#13927)
Co-authored-by: alisonshao <54658187+alisonshao@users.noreply.github.com>
2025-11-25 22:43:19 +08:00
luchangli
64225a8ae9 fix nixl prefill crash make decode health check failed (#13657)
Signed-off-by: liluchang <liluchang@kingsoft.com>
2025-11-25 21:40:11 +08:00
Chen1022
d64bf6c6ce Support piecewise cuda graph for Qwen3-next (#13081) 2025-11-25 21:01:27 +08:00
Liwansi
432ecf841e [Ascend] qwen optimization (#12078)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
2025-11-25 19:44:24 +08:00
Ke Bao
0b3f002daf Update release-whl-kernel.yml (#13921) 2025-11-25 19:05:38 +08:00
Mick
6f094deff0 [diffusion] CI: minor refactor CI for less code duplication (#13905) 2025-11-25 18:44:11 +08:00
ant-yy
59464dbf15 [Fix]: Further fix the buffer len of future map (#13916)
Signed-off-by: vito.yy <vito.yy@antgroup.com>
2025-11-25 18:09:28 +08:00
Yi Zhang
1f7fcc10d5 [diffusion] profile: fix profiling bugs (#13642)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-11-25 17:58:35 +08:00
alisonshao
dbab5d50a3 Add test_dummy_grok_models.py to not_in_ci section (#13908) 2025-11-25 16:51:25 +08:00
Lzhang-hub
760c20b360 update flashinfer_cubin==0.5.3 (#13848) 2025-11-25 00:10:34 -08:00
Xiaoyu Zhang
407cb3ce1e [CI tiny fix] Enhance robustness of vision chunked prefill test with ROUGE-L metric (#13793)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2025-11-25 15:41:14 +08:00
alisonshao
7cc43bd453 Move test_dummy_grok_models.py from manual to srt (temporary) (#13901) 2025-11-24 23:36:53 -08:00
Zaili Wang
cce2d748ef remove RoPE CPU fp32 tests (#13827)
Co-authored-by: Fan Yin <1106310035@qq.com>
2025-11-24 23:22:35 -08:00
DarkSharpness
c1dd9a9599 [Fix] JIT kernel dependencies in other platforms (#13889) 2025-11-24 23:19:17 -08:00
Douglas Yang
ed8786b0b9 Adding nightly tests for Kimi-K2-thinking, Qwen3, minimax-m2, GLM4.6 (#13890) 2025-11-24 22:47:46 -08:00
alisonshao
f9fe06309f Fix trace publish paths in nightly-test-nvidia workflow (#13888) 2025-11-24 21:58:26 -08:00
gongwei-130
8ff3ef1fef fix: draft model revision misuse model revision (#11893) 2025-11-24 21:13:37 -08:00
Yibo Cai
da182e4b83 [CI] fix lint error (#13891) 2025-11-24 20:51:33 -08:00
alisonshao
83e7207763 [diffusion] CI: add validation and cleanup for corrupted safetensors in multimodal loader (#13870)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-11-25 12:02:43 +08:00
Yibo Cai
a2c388ba11 [CI] fix multimodel-gen-test job (#13874) 2025-11-24 19:46:23 -08:00
Mick
9384fa2729 [diffusion] refactor: remove training-related code (#13860) 2025-11-25 11:38:50 +08:00
alisonshao
173e73fa1e Fix nightly test job to fail when any test fails (#13871) 2025-11-24 19:33:46 -08:00
Siyuan Chen
a164259efd [Router bugfix] Fix router_manager selecting the wrong router when enable-igw. (#13572) 2025-11-24 18:51:41 -08:00
Nicolas Castet
b0a26ba624 Add support for bf16 x bf16 cutlass fused MoE (#10275)
Co-authored-by: Sam Li <lsam@nvidia.com>
Co-authored-by: jackeyhua <jackeyhuasjtu@gmail.com>
2025-11-24 18:49:39 -08:00
Binyao Jiang
de430b6745 [Performance] Replace preprocess_video logic from GLM multimodal processor with transformer impl for speed up (up to 27% faster) and addressing OOM (up to 50x improvements) (#13487) 2025-11-24 18:17:13 -08:00
Qiaolin Yu
4b45d556a7 Overlap glm moe gemms in two cuda streams (#13786) 2025-11-24 18:15:24 -08:00
Even Zhou
db0ffc09ef [NPU] Fix NPU CI (#13834)
Co-authored-by: c30031083 <chenxu140@huawei.com>
2025-11-25 10:09:36 +08:00
Glen Liu
eb1d885400 add LoRA warning if loading a preexisting LoRA adapter with a different name (#13822) 2025-11-24 15:16:41 -08:00
Cheng Wan
bf10869203 [Doc] Add an Introduction to Expert Parallelism (#13783) 2025-11-24 14:46:51 -08:00
Lianmin Zheng
e83bd1fadc [Auto Sync] Update schedule_batch.py, schedule_policy.py, b... (20251122) (#13763)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: Hanming Lu <hanming@x.ai>
2025-11-24 14:33:31 -08:00