Commit Graph

6724 Commits

Author SHA1 Message Date
Hubert Lu
36d147121f [AMD] Apply AITER_MXFP4_MOE_SF=1 only to gfx950 in aiter build (#13092) 2025-11-11 12:26:09 -08:00
cctry
e0e6a6efb2 Fix CPP Radix Cache and add test to CI (#11645) 2025-11-11 11:51:37 -08:00
Ziming Huang
e8114102fa [Fix] Update text_chunks in bench_serving chat completions (#13041) 2025-11-11 11:41:11 -08:00
vikram singh shekhawat
14a339fcc8 [Test] Handle streaming chunks with null content in case of stream end. (#10862)
Co-authored-by: svc_repro_tool <svc_repro_tool@habana.ai>
2025-11-12 02:03:50 +08:00
Liangsheng Yin
5f662e786f Revert "[AMD] Add PD test for AMD CI (#11938)" (#13088)
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
2025-11-12 01:23:55 +08:00
Frank Minors
7f5055ed5e Fix cached tokens usage bug (#12814)
Co-authored-by: FrankMinions <liuchen@shinemo.com>
2025-11-12 01:02:43 +08:00
Johnsonms
63728b1154 [Bug] Login shell error: bash: /root/.cargo/env: No such file or directory (#12941) 2025-11-12 00:52:02 +08:00
Ziming Huang
5de25f786e [BugFix] Fix prefill memory leak in PD + GDN (#12994) 2025-11-12 00:48:07 +08:00
Netanel Haber
d52800dbf5 Support file:// scheme in load_video (#13076)
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
2025-11-12 00:36:37 +08:00
yinghui
38a704bccb refine stdout logging codes (#13015) 2025-11-12 00:16:14 +08:00
Edmund Suen
a06c44f905 fix: display served_model_name in /v1/models (#13063) 2025-11-12 00:13:58 +08:00
PiteXChen
527b7d3f59 [RadixTree] Reduce Stack Push/Pop Overhead for Leaf Nodes, Improve radix_tree Leaf Collection Performance (#12199)
Signed-off-by: CLFutureX <chenyongqyl@163.com>
2025-11-11 23:35:39 +08:00
Liangsheng Yin
f09eee036d Tiny simplify evcition metrics collector (#12983) 2025-11-11 23:23:48 +08:00
Ke Bao
e38994dd71 Update rope dtype config (#13037) 2025-11-11 22:47:29 +08:00
yctseng0211
4a78031a71 [ROCM] Optimized deepseek-r1 model with rmsnorm + fp8 quant fusion (#12689)
should be clean after https://github.com/sgl-project/sglang/pull/13017 landed
2025-11-11 02:59:10 -08:00
haoyangli-amd
ea10a9d165 [bug][rocm]fix qr when variable inp (#11609)
Signed-off-by: Haoyang Li <lihaoyang0109@gmail.com>
2025-11-11 01:43:48 -08:00
elvischenv
71aea45c41 [Fix] Add TPOT back to bench_serving (#12976) 2025-11-11 17:32:24 +08:00
Johnsonms
e3b38d7194 [Bug] TypeError: maybe_executor_submit() (#13050) 2025-11-11 00:55:13 -08:00
LHXuuu
5c0cadd0c6 Remove duplicate import (#12980)
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
2025-11-11 16:53:39 +08:00
michael-amd
39b1d048a0 [AMD] Add PD test for AMD CI (#11938) 2025-11-11 16:49:44 +08:00
Yi Zhang
8db7fc4186 disable overlap schedule if mamba radix cache open (#13057) 2025-11-11 16:35:38 +08:00
Xiaoyu Zhang
fe92d4d88e [CI] Auto format code (#13053) 2025-11-10 23:30:29 -08:00
Feng Su
fc8cda14cf Sglang Tracing: optimize trace_event_batch() (#13036)
Signed-off-by: Feng Su <sufeng@linux.alibaba.com>
2025-11-11 15:02:20 +08:00
Mick
6a7322ffbc [diffusion] doc: add support_new_models.md (#13043) 2025-11-11 12:49:48 +08:00
Sai Enduri
c751cb38b0 [AMD CI] Update CI Version Logic. (#13029) 2025-11-10 20:42:00 -08:00
Sam
3594815a8b Re-enable Flashinfer TRTLLM GEN MHA and Add Unit Test (#12885) 2025-11-10 20:17:43 -08:00
Xiaoyu Zhang
9caca6a45c [PieceWise CUDA Graph] Support awq/gptq model in piecewise cudagraph (#12518) 2025-11-11 11:56:15 +08:00
yinghui
08c805a85f fix(ci): workflow id in permission rate limit (#13035) 2025-11-11 11:06:13 +08:00
Xiaoyu Zhang
f18ec927f3 fix tuning_fused_moe_triton_sep tool per_channel_quant bug (#13027) 2025-11-11 10:33:54 +08:00
Sai Enduri
aea88fa7af [AMD CI] Update docker release workflows docker file name. (#13028) 2025-11-10 18:28:39 -08:00
rongfu.leng
2fe4e69fca [router] add postgres databases data connector (#12218) 2025-11-10 16:51:50 -08:00
Peng Zhang
012bfc4fdc [9/n] decouple quantization impl from vllm dependency - adjust ci (#12753) 2025-11-10 14:55:19 -08:00
Keyang Ru
0493775b06 [router][ci] Quick Improvement to make CI more stable (#12869) 2025-11-10 13:56:35 -08:00
Lianmin Zheng
40b26b456b Simplify the BatchMultimodalOutput in io_struct.py (#12993) 2025-11-10 13:54:56 -08:00
Keyang Ru
9840bf4f84 [router][ci] Fix maturin build (#13012) 2025-11-10 12:33:32 -08:00
sglang-bot
303cc957e6 chore: bump SGLang version to 0.5.5.post1 (#13000) 2025-11-10 11:53:43 -08:00
sogalin
661c1c97ad Add pre-suffle weight for new aiter MoE support. (#12908)
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
2025-11-10 11:42:55 -08:00
Kangyan-Zhou
c022107f8b Resolve HF download issue and download models before CI run starts for 8-gpu-h200 runners (#12952) 2025-11-10 11:32:05 -08:00
Liangsheng Yin
56c83e0fb3 [CI] Limit the CI trigger frequency of low-privilege actors (#13010) 2025-11-10 11:27:10 -08:00
Sai Enduri
b51d46d092 [AMD CI] Remove SRT docker build. (#11850)
Co-authored-by: Hubert Lu <Hubert.Lu@amd.com>
2025-11-10 11:25:37 -08:00
Liangsheng Yin
1086473111 Enhance retract test (page cases, long output cases) (#12781) 2025-11-11 03:03:26 +08:00
Liangsheng Yin
665416f6dd Unify memory management across (overlap, non-overlap) x (page>=1) x (spec, non-spec, spec v2) x (retract, finished) (#12224) 2025-11-11 02:56:22 +08:00
Chang Su
838bcb0d93 [misc][ci] Add run-ci after auto-labeler (#13013) 2025-11-10 10:41:54 -08:00
Liangsheng Yin
f1f4c451ab Add process_prefill_chunk back to fix PP event loop (#13009) 2025-11-11 00:51:36 +08:00
fzyzcjy
b0ee99dd03 Super tiny fix typo (#13001) 2025-11-11 00:47:45 +08:00
Mick
ddfcb7c8ab minor: fix notebook bug with new model_info fields added for warmup (#13005) 2025-11-11 00:46:12 +08:00
Ke Bao
58b12ccb46 Support piecewise cuda graph for deepseek v3 (#12996) 2025-11-10 23:18:03 +08:00
Xiaoyu Zhang
547de8c774 [1 / 2] register weak_ref_tensor in sgl-kernel (#12999) 2025-11-10 22:12:59 +08:00
sglang-bot
37c40a87a8 chore: bump sgl-kernel version to 0.3.17 (#12966) 2025-11-10 21:50:58 +08:00
Yuhao Yang
1240ac13b8 vlm: fix tiny multimodal cache bug (#12984) 2025-11-10 21:30:19 +08:00