Commit Graph

2261 Commits

Author SHA1 Message Date
Alison Shao
1aa6ab41de [Nightly] Replace MiniMax-M2 with MiniMax-M2.5 (#20083)
Co-authored-by: Alison Shao <alisonshao@mac.lan>
2026-03-07 01:15:34 -08:00
YC Tseng
c267bdb805 [AMD] Fix AMD CI - stage-b-small-1-gpu-amd (partition 7) (#20028) 2026-03-06 23:49:14 -08:00
Alison Shao
011806c419 [Nightly] Add Kimi K2.5 nightly test (base + Eagle3 MTP), replace Kimi K2 (#19802)
Co-authored-by: Alison Shao <alisonshao@Mac.attlocal.net>
2026-03-06 23:44:04 -08:00
Alison Shao
c584158135 [CI] Temporarily disable flaky test_priority_metrics on CUDA (#20075)
Co-authored-by: Alison Shao <alisonshao@Mac.attlocal.net>
2026-03-06 20:34:45 -08:00
Alison Shao
50bbdcf8e9 Relax flaky test thresholds for MLA DeepSeek V3 and AutoRound (#20068)
Co-authored-by: Alison Shao <alisonshao@Mac.attlocal.net>
2026-03-06 17:26:28 -08:00
shubham singhal
a0d085c16d Adding correct path for module not found error while collecting test (#19778)
Co-authored-by: sys-lpot-val <sys_lpot_val@intel.com>
2026-03-06 16:26:16 -08:00
Alison Shao
ac453b253f Add Qwen3.5-397B-A17B nightly test (8-GPU) (#19906)
Co-authored-by: Alison Shao <alisonshao@mac.lan>
Co-authored-by: Alison Shao <alisonshao@Mac.attlocal.net>
2026-03-06 13:49:28 -08:00
Mohammad Miadh Angkad
8cdb7e1fd4 [CI] Add GPT-OSS test for SM120 (#20056) 2026-03-06 11:43:04 -08:00
Baizhou Zhang
04e364d538 [V32] Enhance deepseek v32 related tests (#19985) 2026-03-05 20:12:49 -08:00
kpham-sgl
346a4131cf [Spec] Refactor NaN/OOB checks to async maybe_detect_* with env-var control (#19899)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-03-05 13:51:05 -08:00
Xinyu Zhang
b3cfad0a80 Add Ray actor support for scheduler process management (DP=1) (#17684)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-03-05 13:21:23 -08:00
sglang-bot
ebb66cc1de [misc] Priority scheduling metrics cleanup (#19927)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-05 12:42:42 -08:00
Michael
203cd8eb02 [AMD] [Z-Image-Turbo Day 0] Add Z-Image-Turbo nightly test for AMD GPUs (#19733) 2026-03-05 08:36:17 -08:00
YC Tseng
b5edab57f2 [AMD] CI - Add MI35x nightly/PR tests for kv-cache-fp8 and allreduce-fusion (DeepSeek) (#19834)
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
2026-03-05 07:09:57 -08:00
Tiwei Bie
727face6c2 [DLLM] Add initial radix cache support (#18724) 2026-03-04 23:24:09 -08:00
Baizhou Zhang
10c65df48a [Bug] Fix lora tp bug on H200 (#19769) 2026-03-04 20:11:02 -08:00
Bruce Changlong Xu
feda2b11c4 [AMD] Add AWQ AMD CI coverage and quantization platform compatibility docs (#19550) 2026-03-04 19:50:55 -08:00
Kangyan-Zhou
198381d9ce Add SSL/TLS support for HTTP and gRPC servers (#18973)
Co-authored-by: guys@spotify.com
2026-03-04 19:27:16 -08:00
Ethan (Yusheng) Su
e555a6c171 [feat] Enhance lora_update_weight_from_tensor for RL training (#19314) 2026-03-04 18:10:42 -08:00
Liangsheng Yin
861d78635f [CI] remove itl testing due to unstable networking (#19904) 2026-03-04 16:32:17 -08:00
zhuxinjie-nz
28c931e1a5 feat: Priority-based scheduling optimization (including default priority, preemption toggle, priority-based metrics, etc.) (#17026)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-03-04 14:52:08 -08:00
hlu1
9457c049e1 [Qwen3.5] Enable MTP spec_v2 and add test for nvidia/Qwen3.5-397B-A17B-NVFP4 (#19391) 2026-03-04 14:01:25 -08:00
Ken J
44208d2adf [vlm][minicpm] support input formats of processor output and embedding (#19614) 2026-03-04 12:11:12 -05:00
strgrb
34c19a32c1 fix flaky test for test_kda_kernels (#19864) 2026-03-04 22:47:29 +08:00
YeChang Guo
6910c1b281 [Feature][NPU]: add runtime support for GPTQ-quantized MoE models (#16364)
Co-authored-by: GuoYechang <52730608+GuoYechang@users.noreply.github.com>
Co-authored-by: root <root@localhost.localdomain>
2026-03-04 16:02:19 +03:00
Charles Chen
d22c6a3847 fix: Properly return abort error for streaming requests if the abort is triggered by scheduler (#19357) 2026-03-03 17:18:15 -08:00
Alison Shao
eb6bcc5c86 [CI] Register test_quant_config_parsing.py in CI suite (#19809)
Co-authored-by: Alison Shao <alisonshao@Mac.attlocal.net>
2026-03-03 16:53:31 -08:00
Hubert Lu
441045a7bf [AMD] Fix EAGLE3 speculative decoding with aiter attention backend (#19362) 2026-03-03 16:12:13 -08:00
Praneth Paruchuri
f7897def96 [Feature] Improve weight loading log (#18651)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-03 14:16:13 -08:00
Guy Stone
f749802402 [Score API][18132] return token usage in Score API response (#18381) 2026-03-03 13:45:35 -08:00
Rahul Vijayaraghavan
ac2819c81f Fix assertion tolerance for bf16 precision in triton attention UT (#17461)
Signed-off-by: Rahul Vijayaraghavan <rahul.vijayaraghavan@intel.com>
2026-03-03 13:43:58 -08:00
Jasonzhang517
d939e26585 [model gateway][0/N] router EPD support: add encoder grpc server backend support (#16552)
Co-authored-by: Zongyao Chen <ZongYao.Chen@linux.alibaba.com>
Co-authored-by: Zongyao Chen <solar1s@163.com>
2026-03-03 19:38:15 +08:00
Muqi Li
666caaf9ce [Tool Call] Stream DeepSeek-V3.2 function call parameters in JSON format. (#16091)
Co-authored-by: Huixxi <uestc.hugo@gmail.com>
2026-03-03 01:46:29 -08:00
Shaun Kotek
4c95953b77 Fix/nemotron mtp quantaized (#19433) 2026-03-03 01:07:46 -08:00
Charles Chen
af0d35b224 Fix: Reject requests with a duplicate request ID which can cause server crash/hang (#19035)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-03-03 00:33:25 -08:00
Zack Yu
07b8d763ef feat: Add FP8 KV cache support for Triton attention backend (#18882) 2026-03-02 23:38:34 -08:00
Xinyuan Tong
dbf1247fe0 Add KimiK2Detector with tool interruption support (#19696)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-03-03 14:04:49 +08:00
Michael
6b8e62f94f [AMD] [Qwen 3.5 Day 0] Add Qwen 3.5 nightly accuracy tests (#19479) 2026-03-02 19:42:42 -08:00
Alison Shao
fe9d85d93c Fix CompressedTensorsMxInt4MoE abstract method and relax GPQA baseline (#19726)
Co-authored-by: Alison Shao <alisonshao@Mac.attlocal.net>
2026-03-02 19:03:21 -08:00
KnightLTC
1041f240c0 [NPU]grok2 model support (#17119)
Co-authored-by: cy <chenyang08056032@163.com>
2026-03-03 10:24:10 +08:00
Glen Liu
cc860a2198 [TestFix] change LoRA tests to use NVIDIA adapter instead of Nutanix (#19642)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-02 12:55:41 -08:00
Yuwei An
c64274c746 Piecewise Cuda Graph set default (#16331) 2026-03-02 23:18:07 +08:00
Leoyzen
da2a0240f7 Add GLM45 tool interruption support (#17714)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-02 19:34:12 +08:00
fzyzcjy
7579ab3f33 Enhance error resilience in dump comparator (#19685) 2026-03-02 19:08:35 +08:00
fzyzcjy
e5ef845cad Support multiple verbosity in dump comparator (#19684) 2026-03-02 18:47:30 +08:00
fzyzcjy
3dd4649b42 Beautify text output in dump comparator (#19683) 2026-03-02 18:47:01 +08:00
fzyzcjy
5bf3deb4bc Trace execution information in dump comparator (#19682) 2026-03-02 18:46:27 +08:00
fzyzcjy
3ebd85bf1c Enhance sglang engine dumping tests in dump comparator (#19681) 2026-03-02 18:46:03 +08:00
fzyzcjy
abdc0ee71f Support directory detection in dump comparator (#19680) 2026-03-02 18:45:35 +08:00
fzyzcjy
6980416149 Support non orthogonal parallel axes and explicit replication annotation in dump comparator (#19679) 2026-03-02 18:44:33 +08:00