Commit Graph

7156 Commits

Author SHA1 Message Date
Minglei Zhu
fafaa2ccea [BugFix] fix outplace_fused_experts missing is_gated (#13864) 2025-11-24 12:36:14 -08:00
Simo Lin
9b4b344115 [model-gateway] add grpc server code owner (#13865) 2025-11-24 12:29:57 -08:00
fzyzcjy
94216a9cc4 Fix quantized moe checker fail for Qwen3 dense fp8 model (#13853) 2025-11-24 11:16:50 -08:00
Yuhao Yao
9535015d05 [Perf] Optimize DeepSeek-R1 w4afp8 glue kernels (#10027)
Co-authored-by: Fan Yin <1106310035@qq.com>
2025-11-24 11:05:38 -08:00
Xinyue Zhang
a3b578fc60 [model-gateway] Refactor router e2e responses tests (#13745)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
2025-11-24 10:58:53 -08:00
Liangsheng Yin
b60e769d0e Tiny unpin uvloop for other backends (#13858) 2025-11-25 01:34:57 +08:00
Zhi Yiliu
a95a38078b [Fix] Fix uvloop get_event_loop() is not suitable for 0.22.x (#13612)
Signed-off-by: lzy <tomlzy213@gmail.com>
Co-authored-by: lzy <tomlzy213@gmail.com>
2025-11-25 01:20:00 +08:00
YAMY
98b38de3f2 Fix: Safe RoPE Cache Expansion to Prevent Position-ID Out-of-Bounds in EAGLE + Long-Sequence Workloads (#11871) 2025-11-25 01:19:06 +08:00
ant-yy
a146f833f1 [Fix]: Adjust FutureMap's token_id_bufs Size to Prevent ChunkedPrefill's next_token_ids from Overwriting Previous Prefill Requests' next_token_id (#13713)
Signed-off-by: vito.yy <vito.yy@antgroup.com>
2025-11-25 01:08:52 +08:00
StonyPort
1dd9a6ae4d Fix TorchAO quant in VLM (#13508)
Co-authored-by: qiuxuan.lzw <qiuxuan.lzw@alibaba-inc.com>
2025-11-24 22:15:40 +08:00
Yuan Luo
8ef11569a2 [VLM] Revise InternVL Piecewise CUDA Graph Supporting (#13846)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-11-24 22:15:10 +08:00
Xiaoyu Zhang
ecefc7904f [sgl-kernel Code Clean] Remove useless lightning_attention kernel (#13819) 2025-11-24 18:26:25 +08:00
gaopengff
aeac622058 [Intel XPU]support xgrammar backend for intel xpu (#13245) 2025-11-24 16:48:00 +08:00
Baizhou Zhang
04b52fa8d6 [chore]Upgrade flashinfer to 0.5.3 (#13751) 2025-11-23 23:38:36 -08:00
Yuhao Yang
e5c0f59133 [diffusion] CI: send nightly-test outputs of diffusion to slack for correctness monitoring (#13833)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-11-24 15:22:26 +08:00
Liangsheng Yin
981ca8313f [misc] Rename minilb install env & remove files & fix lint (#13831) 2025-11-24 13:50:00 +08:00
Mick
414248e0d0 [diffusion] doc: minor update contributing.md with test section (#13792)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-11-24 13:00:08 +08:00
Yuan Luo
f56b9b42e6 [Bugfix] Add jit kernel files in packaging (#13829)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Xu Yongfei <xuyongfei.xyf@antgroup.com>
2025-11-24 12:32:16 +08:00
Liangsheng Yin
b2f7b08c49 Refactor cache init logic (#13800) 2025-11-24 11:41:46 +08:00
Tiance Wang
75222bfed9 Update MindSpore documentation (#13656)
Co-authored-by: wangtiance <tiancew@qq.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-11-24 11:20:51 +08:00
Xiaoyu Zhang
9ea1953331 [Doc] Refine fused_moe_triton configs doc (#13820) 2025-11-23 19:09:41 -08:00
liuhuijiayou
dbf22152d6 Fix bug: Incorrect variable used in rem_total_token_offset calculatio… (#13201) 2025-11-24 11:04:34 +08:00
Swipe4057
d5e0346847 xgrammar up version to 0.1.27 (#13650) 2025-11-24 10:53:45 +08:00
sglang-bot
a22104a676 chore: bump sgl-kernel version to 0.3.18 (#13816) 2025-11-23 17:24:54 -08:00
Baizhou Zhang
4683e244fe [1/2] Refactor DeepGeem requant for FP8 Linear on Blackwell (#13601)
Co-authored-by: fy1214
2025-11-23 16:07:56 -08:00
Simo Lin
9054e844ea remove package json which is not used (#13810) 2025-11-23 15:00:30 -08:00
yinghui
18403f6bfe make trtllm attn backend's init_forward_metadat non blocking (#13802) 2025-11-23 13:35:40 -08:00
Qiaolin Yu
2892265d4c Tune fp8_w8a8 fused triton moe for GLM-4.6-FP8 (#13815) 2025-11-23 13:29:54 -08:00
Baizhou Zhang
c9bd1aca32 [CI] Tiny refactoring sgl-kernel tests (#13813) 2025-11-23 12:45:17 -08:00
hlu1
618ca23802 [Deepseek] Refactor deepseek server_args _handle_model_specific_adjustments (#13687) 2025-11-23 12:41:14 -08:00
Liangsheng Yin
5c2915494c [Scheduler] Tiny organize code style (#13806) 2025-11-23 23:41:34 +08:00
alisonshao
aaa40a9b3b Fix pagination bug in CI monitor preventing performance-test-2-gpu data collection (#13781) 2025-11-23 22:02:30 +08:00
Mick
dd70cf99c1 [diffusion] CI: add run_suite to multimodal_gen CI (#13791) 2025-11-23 21:27:33 +08:00
Mick
d4593964fe [diffusion] feat: support sp for image models (#13180) 2025-11-23 18:11:42 +08:00
YAMY
53fffefd5d Upgrade flashmla kernel for NSA tp support (#13718) 2025-11-23 01:36:49 -08:00
Xiaoyu Zhang
b964ce61d6 [DeepEP] Add SGLANG_DEEPEP_BF16_DISPATCH env var in Normal mode (#13787) 2025-11-23 17:32:43 +08:00
DarkSharpness
ac5505b04c [Feature] HiCache JIT kernel (once again) (#13764) 2025-11-22 22:19:16 -08:00
Peiqi Yin
a90435c059 Fix typo in docs (#13709) 2025-11-23 10:49:49 +08:00
Simo Lin
5354d7b7fd [model-gateway] clean up router manager function order (#13776) 2025-11-22 16:16:20 -08:00
Simo Lin
e01486778a [model-gateway] update smg code owner (#13777) 2025-11-22 15:50:53 -08:00
ErsongWang
047935080d Revert "Fix RMSNorm API CALL mismatch issue. (#10032)" (#13727) 2025-11-22 14:25:52 -08:00
cctry
cad7878964 Gather static input buffers for cuda graph (#13676) 2025-11-22 13:01:50 -08:00
Binyao Jiang
b29769f3b6 Move unnecessary input_addr capture under debug mode flag for speed-up (#13690) 2025-11-22 11:42:26 -08:00
Ho-Ren (Jack) Chuang
3990b84bd3 Refactor MHA & MLA KV caches to support FP4 (#13547)
Signed-off-by: Ho-Ren (Jack) Chuang <horenchuang@bytedance.com>
2025-11-22 11:13:43 -08:00
yinghui
5a4394a342 align code style eagle draft&draft_extend cuda graph runner (#13533) 2025-11-23 01:55:16 +08:00
Yuhao Yang
dd303614e0 [diffusion] CI: tinyfix diffusion ci (#13769)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-11-23 01:03:36 +08:00
Liangsheng Yin
863124684c [Spec v2] Remove allocate_lens and enable over-allocation (#13478) 2025-11-22 22:49:10 +08:00
Yuan Luo
5625e32cae [VLM] Replace torch.repeat_interleave with faster np.repeat for Qwen-VL series (#13736)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-11-22 22:45:32 +08:00
Mick
ca548d8324 [diffusion] refactor: refactor sampling params (#13706) 2025-11-22 22:43:04 +08:00
Liangsheng Yin
3397bcee84 Tiny support different prompts in send_one.py (#13768) 2025-11-22 21:59:58 +08:00