xlzheng
|
37e8724ef3
|
perf: optimize TypeBasedDispatcher using dict for O(1) lookup (#12001)
|
2025-11-15 22:02:42 +08:00 |
|
Zhihao Lyu
|
10592e9c08
|
[Ascend][Feat] Add Ascend sampling backend (#12692)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2025-11-15 21:59:58 +08:00 |
|
Even Zhou
|
2aec8b6e1b
|
[Feature] Spec-Overlap supporting DP-ATTN; PD-Disaggregation; npugraph mode (#12443)
|
2025-11-15 21:51:07 +08:00 |
|
Ke Bao
|
0d41ddfbd0
|
Temporarily disable test_vision_openai_server_a CI (#13331)
|
2025-11-15 19:46:49 +08:00 |
|
fzyzcjy
|
8e6083bfcf
|
Support inverse transform ue8m0 scale (#13285)
|
2025-11-15 16:34:32 +08:00 |
|
alisonshao
|
67e6f1438d
|
Consolidate similar tests to reduce duplication (#12871)
|
2025-11-15 14:29:44 +08:00 |
|
Minglei Zhu
|
a5be6ef98e
|
[Deterministic] Support Qwen3-Next model deterministic inference (#13100)
|
2025-11-14 18:01:18 -08:00 |
|
Kangyan-Zhou
|
b223669136
|
Remove nightly b200 tests and revert a change for test file (#13305)
|
2025-11-14 16:30:22 -08:00 |
|
fzyzcjy
|
15264232ee
|
Super tiny fix CI (#13283)
|
2025-11-14 21:45:10 +08:00 |
|
Sam
|
e7e89349c9
|
Enable Flashinfer TRTLLM-GEN-MoE FP8 blockwise kernel for Qwen3-Next on Blackwell (#12543)
|
2025-11-13 19:44:44 +08:00 |
|
Hubert Lu
|
e4b2937017
|
[AMD] Add AITER Custom All-Reduce (#13102)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2025-11-12 21:53:44 -08:00 |
|
khalilzhk
|
2d531946db
|
[Ascend][feature] support L1+ L2 radixcache on ascend (#12214)
Co-authored-by: khalilzhk <zhangkaikai12@huawei.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
|
2025-11-12 20:45:24 +08:00 |
|
Ke Bao
|
a7cc02e36e
|
Fix run suite sanity check (#13133)
|
2025-11-12 15:29:10 +08:00 |
|
Baizhou Zhang
|
8ece99a9dc
|
[CI] Update job dependency and move dpsk v3.2 tests to 8-gpu suite (#12942)
|
2025-11-11 23:16:47 -08:00 |
|
yctseng0211
|
28b8c5792d
|
Upgrade to ROCm 7.0 image (#13105)
|
2025-11-11 19:14:28 -08:00 |
|
cctry
|
e0e6a6efb2
|
Fix CPP Radix Cache and add test to CI (#11645)
|
2025-11-11 11:51:37 -08:00 |
|
Liangsheng Yin
|
5f662e786f
|
Revert "[AMD] Add PD test for AMD CI (#11938)" (#13088)
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
|
2025-11-12 01:23:55 +08:00 |
|
yctseng0211
|
4a78031a71
|
[ROCM] Optimized deepseek-r1 model with rmsnorm + fp8 quant fusion (#12689)
should be clean after https://github.com/sgl-project/sglang/pull/13017 landed
|
2025-11-11 02:59:10 -08:00 |
|
michael-amd
|
39b1d048a0
|
[AMD] Add PD test for AMD CI (#11938)
|
2025-11-11 16:49:44 +08:00 |
|
Sam
|
3594815a8b
|
Re-enable Flashinfer TRTLLM GEN MHA and Add Unit Test (#12885)
|
2025-11-10 20:17:43 -08:00 |
|
Xiaoyu Zhang
|
9caca6a45c
|
[PieceWise CUDA Graph] Support awq/gptq model in piecewise cudagraph (#12518)
|
2025-11-11 11:56:15 +08:00 |
|
Peng Zhang
|
012bfc4fdc
|
[9/n] decouple quantization impl from vllm dependency - adjust ci (#12753)
|
2025-11-10 14:55:19 -08:00 |
|
Ke Bao
|
db24d34603
|
Support piecewise cuda graph for MLA (#11812)
|
2025-11-10 09:13:48 +08:00 |
|
Liangsheng Yin
|
4f65a64666
|
Refactor / Unify event loop across PD-Disagg, Overlap, DP-Attn cases (#12839)
Co-authored-by: cctry <17473714+cctry@users.noreply.github.com>
|
2025-11-10 00:42:50 +08:00 |
|
Liangsheng Yin
|
877cb52840
|
[CI] increase ut buckets & adjust estimation time. (#12919)
|
2025-11-09 18:17:50 +08:00 |
|
Kangyan-Zhou
|
93cf60fc64
|
Fix Deepseek nightly tests (#12906)
|
2025-11-09 00:26:06 -08:00 |
|
Ke Bao
|
b5e0417392
|
Add kimi k2 thinking to ci (#12907)
|
2025-11-09 16:10:32 +08:00 |
|
Kangyan-Zhou
|
d134096319
|
Add Deepseek models into nightly tests (#12865)
|
2025-11-09 13:17:49 +08:00 |
|
alisonshao
|
d3a03aeef8
|
Refs/heads/add nightly test multi gpu configs (#12870)
|
2025-11-08 15:14:50 -08:00 |
|
Liangsheng Yin
|
6fee2c535c
|
[CI] Tiny adjust CI esitmation time (#12886)
|
2025-11-08 23:02:19 +08:00 |
|
alisonshao
|
0b88d520a0
|
Add nightly performance test for GPT-OSS 4GPU models (#12805)
|
2025-11-07 16:54:07 -08:00 |
|
Ke Bao
|
0fe9c1f70b
|
Fix piecewise cuda graph ci test (#12836)
|
2025-11-08 00:25:00 +08:00 |
|
Jonah Bernard
|
bc25ea6762
|
[MoE] Add Comprehensive MoE Integration Tests (#12090)
|
2025-11-07 00:34:46 -08:00 |
|
Johnsonms
|
125f76ea44
|
[Test] Add DeepSeekV3.2 NSA Indexer Test Suite (#12520)
|
2025-11-06 21:28:43 -08:00 |
|
alisonshao
|
149dc9aab1
|
Add nightly test multi gpu configs (#12721)
|
2025-11-05 19:30:19 -08:00 |
|
Kaixi Hou
|
141278048e
|
[NVIDIA] Fix unit test of MoE and add it to nightly ci (#12709)
|
2025-11-05 14:33:18 -08:00 |
|
Hubert Lu
|
3694266051
|
Expand and update test coverage for AMD CI (#10044)
|
2025-11-04 22:15:13 -08:00 |
|
Kangyan-Zhou
|
6dade6c3b5
|
Fix VLLM dependency test (#12670)
|
2025-11-04 20:49:58 -08:00 |
|
Kaixi Hou
|
0711d1509b
|
[NVIDIA] Fix cutedsl backend of MoE (#12353)
|
2025-11-04 18:54:55 -08:00 |
|
Liangsheng Yin
|
30b26ee9d0
|
Add io struct naming check back (#12634)
|
2025-11-05 01:15:01 +08:00 |
|
Liangsheng Yin
|
aa797d013d
|
[Test] Merge all constrained decoding tests. (#12633)
|
2025-11-05 00:43:06 +08:00 |
|
fzyzcjy
|
b7d7041190
|
Add sanity checks when a test file is not added to CI (reland) (#12594)
|
2025-11-04 18:04:26 +08:00 |
|
Jonah Bernard
|
a209fb05c1
|
[Qwen3 VL] Add LoRA support for Qwen 3 VL (#12165)
|
2025-11-03 20:32:54 -08:00 |
|
Baizhou Zhang
|
9a512cf95b
|
[CI] Move some Lora/Deterministic CI tests to nightly (#12507)
|
2025-11-01 19:54:22 -07:00 |
|
Minglei Zhu
|
229256c505
|
[Deterministic] add deepseek v3 deterministic inference CI test (#12412)
|
2025-11-01 18:10:32 -07:00 |
|
Ke Bao
|
a4bf5c6ad2
|
Support Kimi Linear (#12469)
Co-authored-by: yizhang2077 <1109276519@qq.com>
|
2025-10-31 14:03:35 -07:00 |
|
Neelabh Sinha
|
b57dc169eb
|
[Test] Add Functional Tests for Penalty Parameters (#11931)
|
2025-11-01 00:10:17 +08:00 |
|
Yuhong Guo
|
ab95d35fcb
|
feat: Add Non-intrusive Tensor Dumping for Model Inference (#10566)
|
2025-10-31 12:04:48 +08:00 |
|
Liangsheng Yin
|
d4a09ec9dc
|
[CI] fix tests' time estimation (#12401)
|
2025-10-30 19:20:35 -07:00 |
|
Kaixi Hou
|
c0d02cf4d1
|
[NVIDIA] Add CI workloads for GB200 (#12242)
|
2025-10-30 14:32:03 -07:00 |
|