Yuan Luo
|
ca5c8b16f6
|
[VLM] Support InternVL Vision Encoder Data Parallelism (#13925)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-11-26 11:43:05 +08:00 |
|
timmy-feng
|
f33e5d1ef1
|
fix: spec overlap predict shape does not match verify output shapes (#12786)
|
2025-11-26 11:01:06 +08:00 |
|
Kangyan-Zhou
|
13e5beeab4
|
Fix Deepseek v3.1 loading issue (#13954)
|
2025-11-25 18:28:52 -08:00 |
|
sglang-bot
|
c53e729d45
|
chore: bump sgl-kernel version to 0.3.18.post1 (#13951)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2025-11-25 18:14:28 -08:00 |
|
Kangyan-Zhou
|
03a26557b7
|
Fix nightly-test-nvidia.yml to have the correct trigger (#13950)
|
2025-11-25 17:34:34 -08:00 |
|
Baizhou Zhang
|
873382a910
|
[Tiny]Upgrade README for sgl-kernel (#13945)
|
2025-11-25 16:46:30 -08:00 |
|
sglang-bot
|
391a863b3f
|
chore: bump sgl-kernel version to 0.3.18.post1 (#13942)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2025-11-25 15:31:53 -08:00 |
|
Fan Yin
|
36b1bcd242
|
[chore] update torch version to 2.9 (#12969)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-11-25 14:47:34 -08:00 |
|
fzyzcjy
|
64a11303ce
|
Fix update weight error for blackwell DeepGEMM (#13910)
|
2025-11-25 13:28:12 -08:00 |
|
Lianmin Zheng
|
1ab6ce0e62
|
[Auto Sync] Improve profilers and simplify bench_one_batch_server.py (#13866)
|
2025-11-25 12:13:31 -08:00 |
|
Kangyan-Zhou
|
e99ca6ac74
|
Improve nightly tests (#13903)
|
2025-11-25 12:04:49 -08:00 |
|
Jan Bernlöhr
|
fcccaf9001
|
Add Llama4 attention backend auto-selection (#13421)
Signed-off-by: jbernloehr <jbernloehr@nvidia.com>
|
2025-11-25 11:54:21 -08:00 |
|
Tova Movshovitz
|
215a97fa6c
|
Fix docstrings for v1 HiCacheStorage methods (#13851)
|
2025-11-25 11:48:26 -08:00 |
|
YAMY
|
5eed5fc0b0
|
[DeepSeekV3.2] Centralize NSA dispatch logic in NativeSparseAttnBackend (#13544)
Co-authored-by: hlu1 <14827759+hlu1@users.noreply.github.com>
|
2025-11-25 11:32:30 -08:00 |
|
Baizhou Zhang
|
808b6dfdea
|
[Minor] Fix lint (#13938)
|
2025-11-25 10:57:23 -08:00 |
|
Simo Lin
|
4852aa054c
|
[misc] add llama3.1 chat template (#13935)
|
2025-11-25 09:31:54 -08:00 |
|
Zaili Wang
|
f922bfd520
|
[CPU] Apply PR gating rule in CI workflow (#13933)
|
2025-11-26 00:44:16 +08:00 |
|
Mick
|
dfd7ab9682
|
[diffusion] feat: support LoRA (#13859)
|
2025-11-26 00:21:33 +08:00 |
|
Mick
|
46673b4224
|
[diffusion] doc: add doc for LoRA usage (#13931)
|
2025-11-26 00:02:14 +08:00 |
|
Liangsheng Yin
|
3421d049aa
|
[CI] rename: per_commit -> registered (#13928)
|
2025-11-25 23:05:06 +08:00 |
|
Liangsheng Yin
|
d3d404d3d7
|
[CI] CI registry update (#13927)
Co-authored-by: alisonshao <54658187+alisonshao@users.noreply.github.com>
|
2025-11-25 22:43:19 +08:00 |
|
luchangli
|
64225a8ae9
|
fix nixl prefill crash make decode health check failed (#13657)
Signed-off-by: liluchang <liluchang@kingsoft.com>
|
2025-11-25 21:40:11 +08:00 |
|
Chen1022
|
d64bf6c6ce
|
Support piecewise cuda graph for Qwen3-next (#13081)
|
2025-11-25 21:01:27 +08:00 |
|
Liwansi
|
432ecf841e
|
[Ascend] qwen optimization (#12078)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
|
2025-11-25 19:44:24 +08:00 |
|
Ke Bao
|
0b3f002daf
|
Update release-whl-kernel.yml (#13921)
|
2025-11-25 19:05:38 +08:00 |
|
Mick
|
6f094deff0
|
[diffusion] CI: minor refactor CI for less code duplication (#13905)
|
2025-11-25 18:44:11 +08:00 |
|
ant-yy
|
59464dbf15
|
[Fix]: Further fix the buffer len of future map (#13916)
Signed-off-by: vito.yy <vito.yy@antgroup.com>
|
2025-11-25 18:09:28 +08:00 |
|
Yi Zhang
|
1f7fcc10d5
|
[diffusion] profile: fix profiling bugs (#13642)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-11-25 17:58:35 +08:00 |
|
alisonshao
|
dbab5d50a3
|
Add test_dummy_grok_models.py to not_in_ci section (#13908)
|
2025-11-25 16:51:25 +08:00 |
|
Lzhang-hub
|
760c20b360
|
update flashinfer_cubin==0.5.3 (#13848)
|
2025-11-25 00:10:34 -08:00 |
|
Xiaoyu Zhang
|
407cb3ce1e
|
[CI tiny fix] Enhance robustness of vision chunked prefill test with ROUGE-L metric (#13793)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2025-11-25 15:41:14 +08:00 |
|
alisonshao
|
7cc43bd453
|
Move test_dummy_grok_models.py from manual to srt (temporary) (#13901)
|
2025-11-24 23:36:53 -08:00 |
|
Zaili Wang
|
cce2d748ef
|
remove RoPE CPU fp32 tests (#13827)
Co-authored-by: Fan Yin <1106310035@qq.com>
|
2025-11-24 23:22:35 -08:00 |
|
DarkSharpness
|
c1dd9a9599
|
[Fix] JIT kernel dependencies in other platforms (#13889)
|
2025-11-24 23:19:17 -08:00 |
|
Douglas Yang
|
ed8786b0b9
|
Adding nightly tests for Kimi-K2-thinking, Qwen3, minimax-m2, GLM4.6 (#13890)
|
2025-11-24 22:47:46 -08:00 |
|
alisonshao
|
f9fe06309f
|
Fix trace publish paths in nightly-test-nvidia workflow (#13888)
|
2025-11-24 21:58:26 -08:00 |
|
gongwei-130
|
8ff3ef1fef
|
fix: draft model revision misuse model revision (#11893)
|
2025-11-24 21:13:37 -08:00 |
|
Yibo Cai
|
da182e4b83
|
[CI] fix lint error (#13891)
|
2025-11-24 20:51:33 -08:00 |
|
alisonshao
|
83e7207763
|
[diffusion] CI: add validation and cleanup for corrupted safetensors in multimodal loader (#13870)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-11-25 12:02:43 +08:00 |
|
Yibo Cai
|
a2c388ba11
|
[CI] fix multimodel-gen-test job (#13874)
|
2025-11-24 19:46:23 -08:00 |
|
Mick
|
9384fa2729
|
[diffusion] refactor: remove training-related code (#13860)
|
2025-11-25 11:38:50 +08:00 |
|
alisonshao
|
173e73fa1e
|
Fix nightly test job to fail when any test fails (#13871)
|
2025-11-24 19:33:46 -08:00 |
|
Siyuan Chen
|
a164259efd
|
[Router bugfix] Fix router_manager selecting the wrong router when enable-igw. (#13572)
|
2025-11-24 18:51:41 -08:00 |
|
Nicolas Castet
|
b0a26ba624
|
Add support for bf16 x bf16 cutlass fused MoE (#10275)
Co-authored-by: Sam Li <lsam@nvidia.com>
Co-authored-by: jackeyhua <jackeyhuasjtu@gmail.com>
|
2025-11-24 18:49:39 -08:00 |
|
Binyao Jiang
|
de430b6745
|
[Performance] Replace preprocess_video logic from GLM multimodal processor with transformer impl for speed up (up to 27% faster) and addressing OOM (up to 50x improvements) (#13487)
|
2025-11-24 18:17:13 -08:00 |
|
Qiaolin Yu
|
4b45d556a7
|
Overlap glm moe gemms in two cuda streams (#13786)
|
2025-11-24 18:15:24 -08:00 |
|
Even Zhou
|
db0ffc09ef
|
[NPU] Fix NPU CI (#13834)
Co-authored-by: c30031083 <chenxu140@huawei.com>
|
2025-11-25 10:09:36 +08:00 |
|
Glen Liu
|
eb1d885400
|
add LoRA warning if loading a preexisting LoRA adapter with a different name (#13822)
|
2025-11-24 15:16:41 -08:00 |
|
Cheng Wan
|
bf10869203
|
[Doc] Add an Introduction to Expert Parallelism (#13783)
|
2025-11-24 14:46:51 -08:00 |
|
Lianmin Zheng
|
e83bd1fadc
|
[Auto Sync] Update schedule_batch.py, schedule_policy.py, b... (20251122) (#13763)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: Hanming Lu <hanming@x.ai>
|
2025-11-24 14:33:31 -08:00 |
|