yctseng0211
|
af780c59f3
|
[AMD] add unit-test-backend-8-gpu-amd back (#15253)
Co-authored-by: Sai Enduri <saimanas.enduri@amd.com>
|
2025-12-19 00:59:28 -08:00 |
|
Zehuan Li
|
f6c9db4bc4
|
[DLLM] Fix dLLM regression (#15371)
|
2025-12-19 10:55:46 +08:00 |
|
Zehuan Li
|
b2803ff207
|
[DLLM] Add CI for diffusion LLMs (#14723)
|
2025-12-19 08:54:03 +08:00 |
|
sunxxuns
|
e0963a6cb1
|
[AMD] Clear pre-built AITER kernels and warmup to prevent segfaults and test timeouts (#15318)
|
2025-12-18 14:00:09 -08:00 |
|
Shangming Cai
|
17e81c750a
|
[PP] Add dynamic chunking PP test (#15395)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2025-12-19 01:09:40 +08:00 |
|
Lianmin Zheng
|
d1f0063262
|
Clean up __init__ function of the scheduler and event loop for PD (#15298)
|
2025-12-18 01:35:14 -08:00 |
|
Alison Shao
|
58c840db6b
|
Split test_piecewise_cuda_graph.py to optimize CI resource usage (#15290)
|
2025-12-17 21:20:21 -08:00 |
|
Douglas Yang
|
9e7656be80
|
fix: adjust time for test_epd_disaggregation.py (#15354)
|
2025-12-17 18:30:52 -08:00 |
|
Alison Shao
|
4128d4f5cb
|
[CI] Migrate LoRA tests to test/registered/lora/ (#15176)
|
2025-12-17 13:19:42 -08:00 |
|
Kangyan-Zhou
|
011d8d8970
|
Reserve more memory for DeepSeekOCR model and adjust server start timeout for DeepGEMM to reduce flakiness (#15277)
|
2025-12-17 13:13:05 -08:00 |
|
Liangsheng Yin
|
0c00220795
|
tiny unify environ usage (#15335)
|
2025-12-17 23:31:43 +08:00 |
|
b8zhong
|
ffa7e03506
|
[Piecewise CUDA Graph] Support INT8 (#14918)
|
2025-12-17 18:20:57 +08:00 |
|
Xuchun Shang
|
45a959d3e9
|
[PP] Add pp support for Qwen3-VL (#12333)
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Signed-off-by: Kun(llfl) <i@imux.top>
Signed-off-by: Kun(llfl) <llfl@linux.alibaba.com>
Co-authored-by: kun-llfl <i@imux.top>
Co-authored-by: Kun(llfl) <llfl@linux.alibaba.com>
|
2025-12-17 16:03:58 +08:00 |
|
elvischenv
|
435d1c83c1
|
[Perf] Enable Flashinfer autotune by default (#14357)
|
2025-12-16 23:01:39 -08:00 |
|
Douglas Yang
|
46ad4b986d
|
fix: moving decorator to header (#15297)
|
2025-12-16 17:05:45 -08:00 |
|
Douglas Yang
|
da58df6b3d
|
fix: skipping TestEPDDisaggregationOneEncoder test (#15292)
|
2025-12-16 16:31:20 -08:00 |
|
Alison Shao
|
49237e26fb
|
Fix test_pp_single_node.py estimated time from 800s to 500s (#15291)
|
2025-12-16 16:27:09 -08:00 |
|
amysaq2023
|
ccc8f3b266
|
support non disturbing remote instance weight loader v2 (#14997)
Signed-off-by: Anqi Shen <amy.saq@antgroup.com>
|
2025-12-16 14:39:56 -08:00 |
|
Shangming Cai
|
6682475124
|
[CI] Improve flaky 4 GPU test success rate (#15234)
|
2025-12-17 00:02:26 +08:00 |
|
Liangsheng Yin
|
ecb401ed42
|
Enhance runtime memory check in CI (#15192)
|
2025-12-16 21:40:38 +08:00 |
|
Ke Bao
|
b399e3ac4f
|
Support piecewise cuda graph for fused marlin moe (#15100)
|
2025-12-16 20:05:32 +08:00 |
|
blzheng
|
e27635a02d
|
[CPU] Add 4D input support for ROPE in sgl-kernel (#9337)
|
2025-12-16 17:27:39 +08:00 |
|
Kangyan-Zhou
|
272c5fe43e
|
Increase timeout for TestDeepseekV3MTP for potential DeepGEMM cold start (#15239)
|
2025-12-16 00:26:10 -08:00 |
|
Shangming Cai
|
36fcf71fff
|
[Qwen3-next] Add PD disaggregation support for mamba with extra_buffer (#15180)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com>
|
2025-12-16 14:36:00 +08:00 |
|
XDaoHong
|
4733fcff1f
|
[Feature] npu support enable_torch_compile for torchair backend (#13410)
Co-authored-by: ZhengdQin <zhengdqin@gmail.com>
|
2025-12-16 09:23:51 +08:00 |
|
blzheng
|
d16ff357db
|
[CPU] Add Gemma3RMSNorm kernel in sgl-kernel and add ut (#9324)
|
2025-12-15 00:24:02 -08:00 |
|
Lianmin Zheng
|
702426b06a
|
Fix import warnings (#15144)
|
2025-12-14 21:25:18 -08:00 |
|
Hanming Lu
|
e61dabf5e4
|
[Qwen3-next] support mamba radix cache for overlap scheduler (#14792)
|
2025-12-14 18:54:16 -08:00 |
|
Shangming Cai
|
d277a86dea
|
[CI] Add disaggregation decode PP test (#15114)
|
2025-12-14 23:54:25 +08:00 |
|
Tianyu Guo
|
9acb21ae27
|
feat: support EPD disaggregation (#12263)
Co-authored-by: liusy58 <liusy58@linux.alibaba.com>
Co-authored-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: Nicholas <45984215+liusy58@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
|
2025-12-14 22:30:08 +08:00 |
|
zyl_keep_moving
|
a9ce1623cd
|
[kernel][moe] add moe topk fast (#13969)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
|
2025-12-14 22:26:40 +08:00 |
|
Qiaolin Yu
|
729529190d
|
[ci] Move dpsk-r1-fp4 b200 test to stage b (#15084)
|
2025-12-13 23:14:08 -08:00 |
|
Baizhou Zhang
|
ab3ffd1c8e
|
Add nightly accuracy test for DeepSeek V3.2 (#14935)
|
2025-12-13 12:11:16 -08:00 |
|
Yuhao Yang
|
06b58c5dc5
|
fix flaky image access in ci by switching to raw content url (#14940)
|
2025-12-13 10:52:06 -08:00 |
|
Sam
|
4eda4194f2
|
[Fix] Disable trtllm moe backend for draft model for a qucik fix (#15002)
|
2025-12-12 23:58:47 -08:00 |
|
Baizhou Zhang
|
8698867479
|
[CI]Add gb200 runner back (#15024)
|
2025-12-12 20:19:34 -08:00 |
|
Lianmin Zheng
|
267170bf1d
|
Clean up server args and engine startup processes (#15015)
|
2025-12-12 18:46:07 -08:00 |
|
Yineng Zhang
|
4b7b5af36a
|
Revert several PRs (#14958)
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
|
2025-12-12 11:25:12 -08:00 |
|
Kangyan-Zhou
|
b243154614
|
Fix CI by reverting incorrect metric check logic (#15004)
|
2025-12-12 10:07:38 -08:00 |
|
Yuhao Yao
|
e9e7f15eb5
|
[bugfix] fix TBO crashes when attn_tp_size > 1 (#13730)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2025-12-11 18:18:40 -08:00 |
|
Alison Shao
|
e59435c34b
|
Add retry logic for scheduled CI tests (#14771)
|
2025-12-11 16:59:58 -08:00 |
|
Zaili Wang
|
d6bd2d1126
|
[CPU] layernorm & fused add-layernorm kernels (#14074)
|
2025-12-11 16:58:23 -08:00 |
|
amysaq2023
|
70758d457e
|
support non-disturbing remote-instance-weight-loader (#13125)
Signed-off-by: Anqi Shen <amy.saq@antgroup.com>
|
2025-12-11 16:45:32 -08:00 |
|
Liangsheng Yin
|
543d62d11a
|
Introduce server_fixtures in sglang.test (#14899)
|
2025-12-11 22:30:33 +09:00 |
|
Vladimir221
|
27032cecd9
|
[Ascend]Support of piecewise graph compilation for prefill on NPU (#12287)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2025-12-11 21:10:07 +08:00 |
|
liupeng374
|
388018a5bd
|
[NPU] adapt dsv3.2 nsa prefill context parallel (#14541)
|
2025-12-11 20:12:59 +08:00 |
|
Trang Do
|
8642dbe416
|
Refactor Marlin MoeRunner (#14554)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2025-12-10 19:15:00 -08:00 |
|
Binyao Jiang
|
312df1d6c0
|
Fix TestGLM41VPPAccuracy test flakiness (#14848)
|
2025-12-10 16:59:58 -08:00 |
|
michael-amd
|
c97ce39181
|
[AMD] Add model to AMD nightly test (#14442)
|
2025-12-10 11:59:44 -08:00 |
|
Yuhao Yang
|
c1bd5ee8c5
|
Revert transformers to 4.57.1 (#14801)
|
2025-12-10 11:04:36 -08:00 |
|