Commit Graph

1194 Commits

Author SHA1 Message Date
cctry
e0e6a6efb2 Fix CPP Radix Cache and add test to CI (#11645) 2025-11-11 11:51:37 -08:00
vikram singh shekhawat
14a339fcc8 [Test] Handle streaming chunks with null content in case of stream end. (#10862)
Co-authored-by: svc_repro_tool <svc_repro_tool@habana.ai>
2025-11-12 02:03:50 +08:00
Liangsheng Yin
5f662e786f Revert "[AMD] Add PD test for AMD CI (#11938)" (#13088)
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
2025-11-12 01:23:55 +08:00
Liangsheng Yin
f09eee036d Tiny simplify evcition metrics collector (#12983) 2025-11-11 23:23:48 +08:00
yctseng0211
4a78031a71 [ROCM] Optimized deepseek-r1 model with rmsnorm + fp8 quant fusion (#12689)
should be clean after https://github.com/sgl-project/sglang/pull/13017 landed
2025-11-11 02:59:10 -08:00
haoyangli-amd
ea10a9d165 [bug][rocm]fix qr when variable inp (#11609)
Signed-off-by: Haoyang Li <lihaoyang0109@gmail.com>
2025-11-11 01:43:48 -08:00
michael-amd
39b1d048a0 [AMD] Add PD test for AMD CI (#11938) 2025-11-11 16:49:44 +08:00
Sam
3594815a8b Re-enable Flashinfer TRTLLM GEN MHA and Add Unit Test (#12885) 2025-11-10 20:17:43 -08:00
Xiaoyu Zhang
9caca6a45c [PieceWise CUDA Graph] Support awq/gptq model in piecewise cudagraph (#12518) 2025-11-11 11:56:15 +08:00
Peng Zhang
012bfc4fdc [9/n] decouple quantization impl from vllm dependency - adjust ci (#12753) 2025-11-10 14:55:19 -08:00
Lianmin Zheng
40b26b456b Simplify the BatchMultimodalOutput in io_struct.py (#12993) 2025-11-10 13:54:56 -08:00
Liangsheng Yin
1086473111 Enhance retract test (page cases, long output cases) (#12781) 2025-11-11 03:03:26 +08:00
Liangsheng Yin
665416f6dd Unify memory management across (overlap, non-overlap) x (page>=1) x (spec, non-spec, spec v2) x (retract, finished) (#12224) 2025-11-11 02:56:22 +08:00
fzyzcjy
b0ee99dd03 Super tiny fix typo (#13001) 2025-11-11 00:47:45 +08:00
Yuhao Yang
1240ac13b8 vlm: fix tiny multimodal cache bug (#12984) 2025-11-10 21:30:19 +08:00
Ke Bao
db24d34603 Support piecewise cuda graph for MLA (#11812) 2025-11-10 09:13:48 +08:00
Liangsheng Yin
4f65a64666 Refactor / Unify event loop across PD-Disagg, Overlap, DP-Attn cases (#12839)
Co-authored-by: cctry <17473714+cctry@users.noreply.github.com>
2025-11-10 00:42:50 +08:00
Mick
f5b3ccd9a5 feat: basic support for server-level multimodal cache (#10775) 2025-11-10 00:27:50 +08:00
Ke Bao
bb00e24f87 Adjust server launch time in ci (#12917) 2025-11-09 20:41:29 +08:00
Liangsheng Yin
877cb52840 [CI] increase ut buckets & adjust estimation time. (#12919) 2025-11-09 18:17:50 +08:00
Kangyan-Zhou
93cf60fc64 Fix Deepseek nightly tests (#12906) 2025-11-09 00:26:06 -08:00
Ke Bao
b5e0417392 Add kimi k2 thinking to ci (#12907) 2025-11-09 16:10:32 +08:00
Kangyan-Zhou
d134096319 Add Deepseek models into nightly tests (#12865) 2025-11-09 13:17:49 +08:00
alisonshao
d3a03aeef8 Refs/heads/add nightly test multi gpu configs (#12870) 2025-11-08 15:14:50 -08:00
Liangsheng Yin
6fee2c535c [CI] Tiny adjust CI esitmation time (#12886) 2025-11-08 23:02:19 +08:00
Baizhou Zhang
e039ff382c [CI] Fix huggingface access for test_flash_attention_4.py (#12846) 2025-11-07 20:07:06 -08:00
alisonshao
0b88d520a0 Add nightly performance test for GPT-OSS 4GPU models (#12805) 2025-11-07 16:54:07 -08:00
Ke Bao
0fe9c1f70b Fix piecewise cuda graph ci test (#12836) 2025-11-08 00:25:00 +08:00
Jonah Bernard
bc25ea6762 [MoE] Add Comprehensive MoE Integration Tests (#12090) 2025-11-07 00:34:46 -08:00
Johnsonms
125f76ea44 [Test] Add DeepSeekV3.2 NSA Indexer Test Suite (#12520) 2025-11-06 21:28:43 -08:00
Kalyan Kumar
3b1cc466c0 fixes hardcoded "cuda" device references in unit tests to use a dynamic device selection (#12761) 2025-11-07 11:38:35 +08:00
alisonshao
149dc9aab1 Add nightly test multi gpu configs (#12721) 2025-11-05 19:30:19 -08:00
Lianmin Zheng
c7d57d5bb3 Fix CI and style (#12658) 2025-11-05 15:08:15 -08:00
Kaixi Hou
141278048e [NVIDIA] Fix unit test of MoE and add it to nightly ci (#12709) 2025-11-05 14:33:18 -08:00
Baizhou Zhang
7c45b8b4bb [CI] Fix qwen3-vl lora nightly ci (#12708) 2025-11-05 11:00:13 -08:00
Yuhong Guo
4d84f886e7 Refactor --debug-tensor-dump-layers to list (#12691) 2025-11-05 03:30:01 -08:00
Hubert Lu
3694266051 Expand and update test coverage for AMD CI (#10044) 2025-11-04 22:15:13 -08:00
Kangyan-Zhou
6dade6c3b5 Fix VLLM dependency test (#12670) 2025-11-04 20:49:58 -08:00
Liangsheng Yin
44b1b394a4 [PD-Disagg] Check finish after pop tranferred (#12638) 2025-11-05 11:18:09 +08:00
Kaixi Hou
0711d1509b [NVIDIA] Fix cutedsl backend of MoE (#12353) 2025-11-04 18:54:55 -08:00
soaringk
44da737770 [fix] Handle escaped characters in GLM tool call parser to prevent double serialization (#12456) 2025-11-04 16:48:14 -08:00
Baizhou Zhang
d22d044734 Revert "Enable memory saver for hybrid model" (#12648) 2025-11-04 16:22:06 -08:00
Liangsheng Yin
30b26ee9d0 Add io struct naming check back (#12634) 2025-11-05 01:15:01 +08:00
Liangsheng Yin
aa797d013d [Test] Merge all constrained decoding tests. (#12633) 2025-11-05 00:43:06 +08:00
fzyzcjy
b7d7041190 Add sanity checks when a test file is not added to CI (reland) (#12594) 2025-11-04 18:04:26 +08:00
Minglei Zhu
c14cc47e39 [Deterministic] Optimize bmm_batch_invariant op (#12522) 2025-11-04 00:33:31 -08:00
Zhao Chen
d5fa019c36 feat: limit peak memory usage when computing logprobs (#6318)
Signed-off-by: Zhao Chen <zhaochen.zju@gmail.com>
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
2025-11-03 23:53:20 -08:00
Junrong Lin
173e0f704f Enable memory saver for hybrid model (#11974) 2025-11-04 14:55:26 +08:00
Jonah Bernard
a209fb05c1 [Qwen3 VL] Add LoRA support for Qwen 3 VL (#12165) 2025-11-03 20:32:54 -08:00
kousakawang
7efd8b3d1f [FEAT] Shared mem pool based cuda ipc for multi-modal data transport (#11917)
Co-authored-by: kousakawang <wanghanpei@bytedance.com>
Co-authored-by: Yuan Luo <4908075+yuan-luo@users.noreply.github.com>
2025-11-02 16:46:37 +08:00