Commit Graph

7081 Commits

Author SHA1 Message Date
jacky.cheng
eff7df6d0a [AMD] Enable fused shared expert append and flatten quant for fp8 deepseekR1 model (#13705)
Co-authored-by: yctseng0211 <yctseng@amd.com>
2025-11-21 02:48:28 -08:00
Mick
5e7f91d451 [diffusion] profile: support performance metric dumping and comparison (#13630) 2025-11-21 18:47:16 +08:00
Xiaoyu Zhang
a34d3abb54 [Clean code] Compressed_tensors_moe code clean (#13719) 2025-11-21 18:15:46 +08:00
Cheng Wan
6d0e0b9bfc [11/N] MoE Refactor: Simplifying SBO Implementation with Dispatcher Hooks (#13327)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2025-11-21 01:11:37 -08:00
Even Zhou
589d9ad55b [NPU] chore: bump to CANN 8.3.RC1 and Pytorch 2.8.0 (#13647) 2025-11-21 17:07:08 +08:00
Ziming Huang
a244c0309c [Fix] Qwen3Next lmhead dtype (#13708) 2025-11-21 16:01:33 +08:00
Teng Ma
1bb063aac8 [HiCache] fix unit test with changed new APIs (#13498) 2025-11-21 15:40:31 +08:00
Michele Marzollo
b30f63c40f [Bugfix] Fix hidden state size in EAGLE PD disaggregation buffers (#13590)
Co-authored-by: ZeldaHuang <hzm414167@alibaba-inc.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2025-11-21 15:32:40 +08:00
Lianmin Zheng
90a0133515 [Auto Sync] Update http_server.py, io_struct.py, scheduler_... (20251120) (#13679)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Zhuqi Li <zhli@x.ai>
2025-11-20 23:17:15 -08:00
zyksir
4360279036 [diffusion] server: use meta to avoid Linear init for TextEncoder (#13564)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-11-21 15:14:24 +08:00
ishandhanani
b537ac0d1f Fix ZMQ bind error on non-zero rank nodes when using SGLANG_BLOCK_NONZERO_RANK_CHILDREN=0 (#13686) 2025-11-20 22:54:41 -08:00
Mick
eda2f70033 [diffusion] doc: minor update docs (#13177) 2025-11-21 14:35:29 +08:00
Yuhao Yang
8c212a2029 [difusion] CI: speed up multimodal_gen ci (#13665)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-11-21 14:26:33 +08:00
Stefan He
d754ce973e [Piecewise Cuda Graph] rename, refactor and add more logging (#13675)
Co-authored-by: Minglei Zhu <mingleizhu1122@gmail.com>
Co-authored-by: Ke Bao <ISPObaoke@163.com>
Co-authored-by: Oasis-Git <ayw.sirius19@gmail.com>
2025-11-21 13:28:38 +08:00
Yuan Luo
475962a139 [VLM] Support Piecewise CUDA Graph for InternVL (#13640)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-11-21 13:21:51 +08:00
Xiaoyu Zhang
bfcf15a129 [opt kimi k2 4 / n] Delete useless pad kernel in sgl_moe_align_block_size (#13587) 2025-11-21 13:16:42 +08:00
Xiaoyu Zhang
fb04d43428 [kimi k2 thinking] Avoid useless torch.zeros_ (#13596)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2025-11-21 13:15:27 +08:00
Mike Qiu
6be65ae462 Fix target MLA with eagle3 support for PD disaggregation (#13555)
Signed-off-by: Michael Qiu <qiudayu.qdy@antgroup.com>
Co-authored-by: Mike_Qiu <qiudayu.qdy@antgroup.com>
2025-11-21 12:16:05 +08:00
Glen Liu
750084ae08 remove unnecessary starvation check (#13619) 2025-11-20 19:10:51 -08:00
alisonshao
64480ec712 Add sgl-kernel CI test for Blackwell (B200) (#13301) 2025-11-20 19:02:42 -08:00
Kaixi Hou
db2d362d04 [NVIDIA] Add cutedsl e2e test to GB200 CI (#12672)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2025-11-20 18:30:12 -08:00
alisonshao
81e86992cd [CI] Move nightly tests to test/nightly/ (#13683) 2025-11-20 18:00:02 -08:00
Chang Su
c4db77f8a9 [model-gateway] fix gateway cli arg parser to not use = (#13685) 2025-11-20 17:57:17 -08:00
Mick
c0a2513b07 [diffusion] CI: improve validation method (#13627) 2025-11-21 09:13:13 +08:00
Simo Lin
3ae664d786 [model-gateway] add both python and rust cli alias (#13678) 2025-11-20 17:00:28 -08:00
赵晨阳
c56fc42430 Update quantization.md with new model resources (#13677) 2025-11-20 15:50:16 -08:00
fzyzcjy
3f1cfd87b6 Super tiny remove unused MiniMaxM2MLP class (#13659) 2025-11-21 07:35:12 +08:00
Stefan He
b5344b31b8 [Piecewise CUDA Graph] Fix recompile issue for Mixtral and Grok2 (#13667)
Co-authored-by: Minglei Zhu <mingleizhu1122@gmail.com>
Co-authored-by: Ke Bao <ISPObaoke@163.com>
Co-authored-by: Oasis-Git <ayw.sirius19@gmail.com>
2025-11-20 14:20:11 -08:00
alisonshao
6b262ac839 Test reorganization: Move tests to manual/ (#13610) 2025-11-20 13:41:58 -08:00
Glen Liu
ada8ce1fd0 allow loras to be implicitly evicted and loaded based on max_loaded_loras (#11526) 2025-11-20 13:34:32 -08:00
alisonshao
5a2c70396e Add nightly test CI monitor workflow (#13038) 2025-11-20 12:55:13 -08:00
b8zhong
42028af614 enable csgmv automatically on cuda (#13600) 2025-11-20 12:53:02 -08:00
hlu1
7291c72e57 [Deepseek V3.2] Change indexer weights_proj to fp32 (#13459) 2025-11-20 12:24:10 -08:00
Vincent Zhong
6bc3062894 Fix launch of Olmo3 (#13666)
Signed-off-by: Vincent Zhong <207368749+vincentzed@users.noreply.github.com>
2025-11-20 12:04:30 -08:00
YAMY
fa92441027 [DeepseekV3.2] Deepseek fp8 support for MHA path (#12964) 2025-11-20 11:13:36 -08:00
Douglas Yang
fc9efdcb98 Adding nightly tests as release guard for bot bump workflows (#13655) 2025-11-20 10:04:36 -08:00
Fan Yin
2dec555d36 [10/n] decouple quantization impl from vllm dependency - fix import (#13524) 2025-11-21 01:45:51 +08:00
YanbingJiang
acde21d8d5 Add fused_rmsnorm_gated_cpu kernel for CPU to support Qwen3-Next (#11577) 2025-11-21 01:33:31 +08:00
Liangsheng Yin
4528cb7d40 [CI] apply pr-gate for XPU (#13663) 2025-11-20 23:07:04 +08:00
Lianmin Zheng
a352e833c4 CI: Kill zombie diffusion processes in CI & minor code style fix on rotary embedding fallback (#13637) 2025-11-20 21:57:01 +08:00
Liangsheng Yin
852eb6ce2a [CI] optimize CI workflow info (#13634) 2025-11-20 18:12:51 +08:00
Lianmin Zheng
7af9b88c6c Revert "[Feature] Introduce JIT Kernel in sglang (with hicache JIT kernel)" (#13644) 2025-11-20 02:11:12 -08:00
Lzhang-hub
2847e5c4b4 fix bench_speculative bug (#13197) 2025-11-20 17:09:04 +08:00
yctseng0211
c8ede0e93c [ROCM] Optimized deepseek-r1 fp8 model with + triton_gemm_a8w8 + batch_gemm_a8w8 + fused set_mla_kv_buffer kernel (#13617)
Co-authored-by: root <root@smci355-ccs-aus-m12-17.cs-aus.dcgpu>
Co-authored-by: jacky.cheng <yichiche@amd.com>
2025-11-20 00:29:56 -08:00
Liangsheng Yin
19729f723e [CI] Align metric units for CI rate limit (#13633)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-11-20 16:25:57 +08:00
DarkSharpness
b51f9bbee7 [Feature] Introduce JIT Kernel in sglang (with hicache JIT kernel) (#13453) 2025-11-20 00:03:32 -08:00
Mick
4a8442af1b diffusion: improve baseline performance monitor (#13614) 2025-11-20 14:29:56 +08:00
Thomas Wang
7dcf910d53 Add support for new aiter version (AR accuracy, is_shuffled PR) (#13554)
Co-authored-by: sogalin <39478626+sogalin@users.noreply.github.com>
2025-11-19 22:17:57 -08:00
joesun
c7b37b7074 [diffusion] fix: remove multimodal_gen redundant get_bool_env_var func (#13583)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-11-20 14:05:29 +08:00
iLeGend
10e0b83a4c Add FP32 dtype support for RoPE - Part2 (#13328) 2025-11-19 21:19:53 -08:00