Hubert Lu
|
36d147121f
|
[AMD] Apply AITER_MXFP4_MOE_SF=1 only to gfx950 in aiter build (#13092)
|
2025-11-11 12:26:09 -08:00 |
|
cctry
|
e0e6a6efb2
|
Fix CPP Radix Cache and add test to CI (#11645)
|
2025-11-11 11:51:37 -08:00 |
|
Ziming Huang
|
e8114102fa
|
[Fix] Update text_chunks in bench_serving chat completions (#13041)
|
2025-11-11 11:41:11 -08:00 |
|
vikram singh shekhawat
|
14a339fcc8
|
[Test] Handle streaming chunks with null content in case of stream end. (#10862)
Co-authored-by: svc_repro_tool <svc_repro_tool@habana.ai>
|
2025-11-12 02:03:50 +08:00 |
|
Liangsheng Yin
|
5f662e786f
|
Revert "[AMD] Add PD test for AMD CI (#11938)" (#13088)
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
|
2025-11-12 01:23:55 +08:00 |
|
Frank Minors
|
7f5055ed5e
|
Fix cached tokens usage bug (#12814)
Co-authored-by: FrankMinions <liuchen@shinemo.com>
|
2025-11-12 01:02:43 +08:00 |
|
Johnsonms
|
63728b1154
|
[Bug] Login shell error: bash: /root/.cargo/env: No such file or directory (#12941)
|
2025-11-12 00:52:02 +08:00 |
|
Ziming Huang
|
5de25f786e
|
[BugFix] Fix prefill memory leak in PD + GDN (#12994)
|
2025-11-12 00:48:07 +08:00 |
|
Netanel Haber
|
d52800dbf5
|
Support file:// scheme in load_video (#13076)
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
|
2025-11-12 00:36:37 +08:00 |
|
yinghui
|
38a704bccb
|
refine stdout logging codes (#13015)
|
2025-11-12 00:16:14 +08:00 |
|
Edmund Suen
|
a06c44f905
|
fix: display served_model_name in /v1/models (#13063)
|
2025-11-12 00:13:58 +08:00 |
|
PiteXChen
|
527b7d3f59
|
[RadixTree] Reduce Stack Push/Pop Overhead for Leaf Nodes, Improve radix_tree Leaf Collection Performance (#12199)
Signed-off-by: CLFutureX <chenyongqyl@163.com>
|
2025-11-11 23:35:39 +08:00 |
|
Liangsheng Yin
|
f09eee036d
|
Tiny simplify evcition metrics collector (#12983)
|
2025-11-11 23:23:48 +08:00 |
|
Ke Bao
|
e38994dd71
|
Update rope dtype config (#13037)
|
2025-11-11 22:47:29 +08:00 |
|
yctseng0211
|
4a78031a71
|
[ROCM] Optimized deepseek-r1 model with rmsnorm + fp8 quant fusion (#12689)
should be clean after https://github.com/sgl-project/sglang/pull/13017 landed
|
2025-11-11 02:59:10 -08:00 |
|
haoyangli-amd
|
ea10a9d165
|
[bug][rocm]fix qr when variable inp (#11609)
Signed-off-by: Haoyang Li <lihaoyang0109@gmail.com>
|
2025-11-11 01:43:48 -08:00 |
|
elvischenv
|
71aea45c41
|
[Fix] Add TPOT back to bench_serving (#12976)
|
2025-11-11 17:32:24 +08:00 |
|
Johnsonms
|
e3b38d7194
|
[Bug] TypeError: maybe_executor_submit() (#13050)
|
2025-11-11 00:55:13 -08:00 |
|
LHXuuu
|
5c0cadd0c6
|
Remove duplicate import (#12980)
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
|
2025-11-11 16:53:39 +08:00 |
|
michael-amd
|
39b1d048a0
|
[AMD] Add PD test for AMD CI (#11938)
|
2025-11-11 16:49:44 +08:00 |
|
Yi Zhang
|
8db7fc4186
|
disable overlap schedule if mamba radix cache open (#13057)
|
2025-11-11 16:35:38 +08:00 |
|
Xiaoyu Zhang
|
fe92d4d88e
|
[CI] Auto format code (#13053)
|
2025-11-10 23:30:29 -08:00 |
|
Feng Su
|
fc8cda14cf
|
Sglang Tracing: optimize trace_event_batch() (#13036)
Signed-off-by: Feng Su <sufeng@linux.alibaba.com>
|
2025-11-11 15:02:20 +08:00 |
|
Mick
|
6a7322ffbc
|
[diffusion] doc: add support_new_models.md (#13043)
|
2025-11-11 12:49:48 +08:00 |
|
Sai Enduri
|
c751cb38b0
|
[AMD CI] Update CI Version Logic. (#13029)
|
2025-11-10 20:42:00 -08:00 |
|
Sam
|
3594815a8b
|
Re-enable Flashinfer TRTLLM GEN MHA and Add Unit Test (#12885)
|
2025-11-10 20:17:43 -08:00 |
|
Xiaoyu Zhang
|
9caca6a45c
|
[PieceWise CUDA Graph] Support awq/gptq model in piecewise cudagraph (#12518)
|
2025-11-11 11:56:15 +08:00 |
|
yinghui
|
08c805a85f
|
fix(ci): workflow id in permission rate limit (#13035)
|
2025-11-11 11:06:13 +08:00 |
|
Xiaoyu Zhang
|
f18ec927f3
|
fix tuning_fused_moe_triton_sep tool per_channel_quant bug (#13027)
|
2025-11-11 10:33:54 +08:00 |
|
Sai Enduri
|
aea88fa7af
|
[AMD CI] Update docker release workflows docker file name. (#13028)
|
2025-11-10 18:28:39 -08:00 |
|
rongfu.leng
|
2fe4e69fca
|
[router] add postgres databases data connector (#12218)
|
2025-11-10 16:51:50 -08:00 |
|
Peng Zhang
|
012bfc4fdc
|
[9/n] decouple quantization impl from vllm dependency - adjust ci (#12753)
|
2025-11-10 14:55:19 -08:00 |
|
Keyang Ru
|
0493775b06
|
[router][ci] Quick Improvement to make CI more stable (#12869)
|
2025-11-10 13:56:35 -08:00 |
|
Lianmin Zheng
|
40b26b456b
|
Simplify the BatchMultimodalOutput in io_struct.py (#12993)
|
2025-11-10 13:54:56 -08:00 |
|
Keyang Ru
|
9840bf4f84
|
[router][ci] Fix maturin build (#13012)
|
2025-11-10 12:33:32 -08:00 |
|
sglang-bot
|
303cc957e6
|
chore: bump SGLang version to 0.5.5.post1 (#13000)
|
2025-11-10 11:53:43 -08:00 |
|
sogalin
|
661c1c97ad
|
Add pre-suffle weight for new aiter MoE support. (#12908)
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
|
2025-11-10 11:42:55 -08:00 |
|
Kangyan-Zhou
|
c022107f8b
|
Resolve HF download issue and download models before CI run starts for 8-gpu-h200 runners (#12952)
|
2025-11-10 11:32:05 -08:00 |
|
Liangsheng Yin
|
56c83e0fb3
|
[CI] Limit the CI trigger frequency of low-privilege actors (#13010)
|
2025-11-10 11:27:10 -08:00 |
|
Sai Enduri
|
b51d46d092
|
[AMD CI] Remove SRT docker build. (#11850)
Co-authored-by: Hubert Lu <Hubert.Lu@amd.com>
|
2025-11-10 11:25:37 -08:00 |
|
Liangsheng Yin
|
1086473111
|
Enhance retract test (page cases, long output cases) (#12781)
|
2025-11-11 03:03:26 +08:00 |
|
Liangsheng Yin
|
665416f6dd
|
Unify memory management across (overlap, non-overlap) x (page>=1) x (spec, non-spec, spec v2) x (retract, finished) (#12224)
|
2025-11-11 02:56:22 +08:00 |
|
Chang Su
|
838bcb0d93
|
[misc][ci] Add run-ci after auto-labeler (#13013)
|
2025-11-10 10:41:54 -08:00 |
|
Liangsheng Yin
|
f1f4c451ab
|
Add process_prefill_chunk back to fix PP event loop (#13009)
|
2025-11-11 00:51:36 +08:00 |
|
fzyzcjy
|
b0ee99dd03
|
Super tiny fix typo (#13001)
|
2025-11-11 00:47:45 +08:00 |
|
Mick
|
ddfcb7c8ab
|
minor: fix notebook bug with new model_info fields added for warmup (#13005)
|
2025-11-11 00:46:12 +08:00 |
|
Ke Bao
|
58b12ccb46
|
Support piecewise cuda graph for deepseek v3 (#12996)
|
2025-11-10 23:18:03 +08:00 |
|
Xiaoyu Zhang
|
547de8c774
|
[1 / 2] register weak_ref_tensor in sgl-kernel (#12999)
|
2025-11-10 22:12:59 +08:00 |
|
sglang-bot
|
37c40a87a8
|
chore: bump sgl-kernel version to 0.3.17 (#12966)
|
2025-11-10 21:50:58 +08:00 |
|
Yuhao Yang
|
1240ac13b8
|
vlm: fix tiny multimodal cache bug (#12984)
|
2025-11-10 21:30:19 +08:00 |
|