Zehuan Li
|
21b0582d4b
|
[feature] Initial block diffusion language model support (#12588)
Co-authored-by: Tiwei Bie <tiwei.btw@antgroup.com>
|
2025-11-26 17:57:54 +08:00 |
|
Yuhao Yang
|
5795da5e83
|
[diffusion] fix: fix the issue where the qwen-edit & wan model produces incorrect output during sequence parallelism (#13922)
|
2025-11-26 15:58:23 +08:00 |
|
StonyPort
|
540d6fee20
|
Support piecewise CUDA graph for embedding models (#13852)
Co-authored-by: qiuxuan.lzw <qiuxuan.lzw@alibaba-inc.com>
|
2025-11-26 15:29:46 +08:00 |
|
Thomas Wang
|
5a8adca900
|
Turn off PREBUILD aiter in MI355 (#13963)
|
2025-11-25 22:30:39 -08:00 |
|
ShawnY112358
|
007c3e234c
|
[feat] support in-flight weight update (#10071)
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
|
2025-11-25 22:03:13 -08:00 |
|
fzyzcjy
|
7130ad3a29
|
Fix SGLANG_ENABLE_HEALTH_ENDPOINT_GENERATION not working (#13961)
|
2025-11-25 21:58:55 -08:00 |
|
Zaili Wang
|
35a4c21a8a
|
update CI permission list (#13962)
|
2025-11-25 20:16:44 -08:00 |
|
alisonshao
|
846ba3c62c
|
Fix nightly test failures: NSA indexer dtype and CPP radix cache init (#13958)
|
2025-11-26 11:52:08 +08:00 |
|
Yubo Wang
|
18fb51583f
|
Support FlashAttention3 page_size > 1 and topk > 1 case with paged attn and spec decode (#7725)
|
2025-11-26 11:44:41 +08:00 |
|
Yuan Luo
|
ca5c8b16f6
|
[VLM] Support InternVL Vision Encoder Data Parallelism (#13925)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-11-26 11:43:05 +08:00 |
|
timmy-feng
|
f33e5d1ef1
|
fix: spec overlap predict shape does not match verify output shapes (#12786)
|
2025-11-26 11:01:06 +08:00 |
|
Kangyan-Zhou
|
13e5beeab4
|
Fix Deepseek v3.1 loading issue (#13954)
|
2025-11-25 18:28:52 -08:00 |
|
sglang-bot
|
c53e729d45
|
chore: bump sgl-kernel version to 0.3.18.post1 (#13951)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2025-11-25 18:14:28 -08:00 |
|
Kangyan-Zhou
|
03a26557b7
|
Fix nightly-test-nvidia.yml to have the correct trigger (#13950)
|
2025-11-25 17:34:34 -08:00 |
|
Baizhou Zhang
|
873382a910
|
[Tiny]Upgrade README for sgl-kernel (#13945)
|
2025-11-25 16:46:30 -08:00 |
|
sglang-bot
|
391a863b3f
|
chore: bump sgl-kernel version to 0.3.18.post1 (#13942)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2025-11-25 15:31:53 -08:00 |
|
Fan Yin
|
36b1bcd242
|
[chore] update torch version to 2.9 (#12969)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-11-25 14:47:34 -08:00 |
|
fzyzcjy
|
64a11303ce
|
Fix update weight error for blackwell DeepGEMM (#13910)
|
2025-11-25 13:28:12 -08:00 |
|
Lianmin Zheng
|
1ab6ce0e62
|
[Auto Sync] Improve profilers and simplify bench_one_batch_server.py (#13866)
|
2025-11-25 12:13:31 -08:00 |
|
Kangyan-Zhou
|
e99ca6ac74
|
Improve nightly tests (#13903)
|
2025-11-25 12:04:49 -08:00 |
|
Jan Bernlöhr
|
fcccaf9001
|
Add Llama4 attention backend auto-selection (#13421)
Signed-off-by: jbernloehr <jbernloehr@nvidia.com>
|
2025-11-25 11:54:21 -08:00 |
|
Tova Movshovitz
|
215a97fa6c
|
Fix docstrings for v1 HiCacheStorage methods (#13851)
|
2025-11-25 11:48:26 -08:00 |
|
YAMY
|
5eed5fc0b0
|
[DeepSeekV3.2] Centralize NSA dispatch logic in NativeSparseAttnBackend (#13544)
Co-authored-by: hlu1 <14827759+hlu1@users.noreply.github.com>
|
2025-11-25 11:32:30 -08:00 |
|
Baizhou Zhang
|
808b6dfdea
|
[Minor] Fix lint (#13938)
|
2025-11-25 10:57:23 -08:00 |
|
Simo Lin
|
4852aa054c
|
[misc] add llama3.1 chat template (#13935)
|
2025-11-25 09:31:54 -08:00 |
|
Zaili Wang
|
f922bfd520
|
[CPU] Apply PR gating rule in CI workflow (#13933)
|
2025-11-26 00:44:16 +08:00 |
|
Mick
|
dfd7ab9682
|
[diffusion] feat: support LoRA (#13859)
|
2025-11-26 00:21:33 +08:00 |
|
Mick
|
46673b4224
|
[diffusion] doc: add doc for LoRA usage (#13931)
|
2025-11-26 00:02:14 +08:00 |
|
Liangsheng Yin
|
3421d049aa
|
[CI] rename: per_commit -> registered (#13928)
|
2025-11-25 23:05:06 +08:00 |
|
Liangsheng Yin
|
d3d404d3d7
|
[CI] CI registry update (#13927)
Co-authored-by: alisonshao <54658187+alisonshao@users.noreply.github.com>
|
2025-11-25 22:43:19 +08:00 |
|
luchangli
|
64225a8ae9
|
fix nixl prefill crash make decode health check failed (#13657)
Signed-off-by: liluchang <liluchang@kingsoft.com>
|
2025-11-25 21:40:11 +08:00 |
|
Chen1022
|
d64bf6c6ce
|
Support piecewise cuda graph for Qwen3-next (#13081)
|
2025-11-25 21:01:27 +08:00 |
|
Liwansi
|
432ecf841e
|
[Ascend] qwen optimization (#12078)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
|
2025-11-25 19:44:24 +08:00 |
|
Ke Bao
|
0b3f002daf
|
Update release-whl-kernel.yml (#13921)
|
2025-11-25 19:05:38 +08:00 |
|
Mick
|
6f094deff0
|
[diffusion] CI: minor refactor CI for less code duplication (#13905)
|
2025-11-25 18:44:11 +08:00 |
|
ant-yy
|
59464dbf15
|
[Fix]: Further fix the buffer len of future map (#13916)
Signed-off-by: vito.yy <vito.yy@antgroup.com>
|
2025-11-25 18:09:28 +08:00 |
|
Yi Zhang
|
1f7fcc10d5
|
[diffusion] profile: fix profiling bugs (#13642)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-11-25 17:58:35 +08:00 |
|
alisonshao
|
dbab5d50a3
|
Add test_dummy_grok_models.py to not_in_ci section (#13908)
|
2025-11-25 16:51:25 +08:00 |
|
Lzhang-hub
|
760c20b360
|
update flashinfer_cubin==0.5.3 (#13848)
|
2025-11-25 00:10:34 -08:00 |
|
Xiaoyu Zhang
|
407cb3ce1e
|
[CI tiny fix] Enhance robustness of vision chunked prefill test with ROUGE-L metric (#13793)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2025-11-25 15:41:14 +08:00 |
|
alisonshao
|
7cc43bd453
|
Move test_dummy_grok_models.py from manual to srt (temporary) (#13901)
|
2025-11-24 23:36:53 -08:00 |
|
Zaili Wang
|
cce2d748ef
|
remove RoPE CPU fp32 tests (#13827)
Co-authored-by: Fan Yin <1106310035@qq.com>
|
2025-11-24 23:22:35 -08:00 |
|
DarkSharpness
|
c1dd9a9599
|
[Fix] JIT kernel dependencies in other platforms (#13889)
|
2025-11-24 23:19:17 -08:00 |
|
Douglas Yang
|
ed8786b0b9
|
Adding nightly tests for Kimi-K2-thinking, Qwen3, minimax-m2, GLM4.6 (#13890)
|
2025-11-24 22:47:46 -08:00 |
|
alisonshao
|
f9fe06309f
|
Fix trace publish paths in nightly-test-nvidia workflow (#13888)
|
2025-11-24 21:58:26 -08:00 |
|
gongwei-130
|
8ff3ef1fef
|
fix: draft model revision misuse model revision (#11893)
|
2025-11-24 21:13:37 -08:00 |
|
Yibo Cai
|
da182e4b83
|
[CI] fix lint error (#13891)
|
2025-11-24 20:51:33 -08:00 |
|
alisonshao
|
83e7207763
|
[diffusion] CI: add validation and cleanup for corrupted safetensors in multimodal loader (#13870)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-11-25 12:02:43 +08:00 |
|
Yibo Cai
|
a2c388ba11
|
[CI] fix multimodel-gen-test job (#13874)
|
2025-11-24 19:46:23 -08:00 |
|
Mick
|
9384fa2729
|
[diffusion] refactor: remove training-related code (#13860)
|
2025-11-25 11:38:50 +08:00 |
|