Commit Graph

6776 Commits

Author SHA1 Message Date
Vladimir Serov
c4b74c1db2 [Ascend] LoRA: adding Ascend LoRA backend with using kernels from sgl_kernel_npu (#12288)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>
2025-11-12 16:20:34 -08:00
Shu Wang
7aa443903d Fix nan in global scaling factor for large scale nvfp4 EP (#13162) 2025-11-12 15:32:23 -08:00
bingps
4eda9969e8 [DeepseekV32]: use _concat_mla_absorb_q_general to replace torch.cat (#12215)
Co-authored-by: Guangda Liu <bingps@users.noreply.github.com>
2025-11-12 15:20:01 -08:00
Douglas Yang
03a7e6f4db Add job and runner failure monitor workflow for CI (#13104) 2025-11-12 14:04:50 -08:00
Xinyue Zhang
2cdde3d46d [router] Fix Flaky test_circuit_breaker_opens_and_recovers (#13164) 2025-11-12 12:07:42 -08:00
Keyang Ru
401ed0c594 [router] Add comprehensive validation to Responses API (#13127) 2025-11-12 11:19:27 -08:00
Siyuan Chen
4ef4390540 bugfix: multi-model routing for /generate api (#12979)
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
Co-authored-by: Chang Su <chang.s.su@oracle.com>
2025-11-12 11:13:27 -08:00
Lianmin Zheng
d646cf6347 [Auto Sync] Update test_deterministic.py (20251112) (#13128)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
2025-11-12 11:12:23 -08:00
Ke Bao
4edb240112 Fuse routed_scaling_factor to fused_marlin_moe (#12998) 2025-11-13 00:18:58 +08:00
Yuan Luo
706502ff6c [VLM] Support PP for Qwen2.5-VL (#13075)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn>
2025-11-12 23:18:44 +08:00
Makcum888e
c2e56dadb2 [Ascend] torch_npu.npu_mrope for MRotaryEmbedding (#10907)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>
2025-11-12 21:45:54 +08:00
Zhihao Zhang
9c1c5c6d7d [ngram] use SGLANG_NGRAM_FORCE_GREEDY_VERIFY to control verify method (#13153)
Co-authored-by: a4zhangfei <a4zhangfei@qq.com>
2025-11-12 21:07:33 +08:00
khalilzhk
2d531946db [Ascend][feature] support L1+ L2 radixcache on ascend (#12214)
Co-authored-by: khalilzhk <zhangkaikai12@huawei.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
2025-11-12 20:45:24 +08:00
sglang-bot
ebaf86d441 chore: bump SGLang version to 0.5.5.post2 (#13129)
Include the critical fix https://github.com/sgl-project/sglang/pull/12915.
2025-11-12 20:35:20 +08:00
zhanghaotong
8359f185a1 [Feature] Propagate Trace Headers into Root Span for OpenTelemetry Cross-Service Context (#10808)
Signed-off-by: zhanghaotong <zhanghaotong.zht@antgroup.com>
2025-11-12 20:28:59 +08:00
ronnie_zheng
5324f37ab3 [Ascend]adapt enable-profile-cuda-graph for NPU (#12617)
Co-authored-by: Leo920320 <1530593013@qq.com>
Co-authored-by: liupengcheng <6531487+leo920320@user.noreply.gitee.com>
2025-11-12 19:57:24 +08:00
Charlie Ruan
2864c49fbc [RPC] Fix handle_rpc_request with **recv_req.parameters (#7906) 2025-11-12 19:29:12 +08:00
sogalin
d26ec39ff2 Update aiter to v0.1.7.post1 (#13149) 2025-11-12 03:09:44 -08:00
Yikai Zhang
9c546bfd6d Fix the Wrong Return Type of Scheduler.recv_requests (#7886) 2025-11-12 18:54:56 +08:00
Rohan Potdar
c9b581644d Dump total_throughput to output-file in bench_serving.py (#9790) 2025-11-12 18:48:28 +08:00
Chang Su
e5e65e3d2a [router][grpc] Support vllm backend for grpc router (#13120) 2025-11-12 02:29:20 -08:00
SijiaYang
ffeb28ba6f fix: duplicate resize images logic of qwen-vl series models (#12458)
Signed-off-by: yangsijia.614 <yangsijia.614@bytedance.com>
2025-11-12 18:08:40 +08:00
Jimmy
b40f605fde fix(tcp-port): replace bind_server_socket to get_zmq_socket(Port conflict) (#11961)
Co-authored-by: wangchao <wcsjtu@163.com>
2025-11-12 18:05:16 +08:00
Simo Lin
3cdec20c6b [router] add minmax m2 reasoning parser (#13137) 2025-11-12 18:27:05 +09:00
Danylo Vashchilenko
d28caaf60a [router] Support complex assistant and tool messages in /chat/completions (#12860)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
2025-11-12 00:14:15 -08:00
Liangsheng Yin
ad8d24c39e Fix re-trigger actor of CI rate limit (#13136) 2025-11-12 16:10:27 +08:00
Chang Su
018123b5b0 [misc] Remove performance and router-benchmark label matching (#13135) 2025-11-12 00:04:20 -08:00
Xinyuan Tong
4983b7e7aa Fix strict level setting for Kimi K2 tool calls when not explicitly set (#13077)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2025-11-12 00:01:42 -08:00
wxsm
f825137f8a fix: Remove dupulicated kv_events initialization in scheduler (#13132) 2025-11-11 23:54:47 -08:00
Simo Lin
ae68158f2e [router] move radix tree to policy crate and addreses some code styles (#13131) 2025-11-11 23:36:04 -08:00
Ke Bao
a7cc02e36e Fix run suite sanity check (#13133) 2025-11-12 15:29:10 +08:00
Baizhou Zhang
8ece99a9dc [CI] Update job dependency and move dpsk v3.2 tests to 8-gpu suite (#12942) 2025-11-11 23:16:47 -08:00
Teng Ma
44e391b6e9 [PD] Add custom gpu id to device topo support (#12817)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2025-11-12 15:00:45 +08:00
ishandhanani
1a5c313f97 feat(engine): add rid parameter to methods in Engine class (#13095) 2025-11-11 22:37:54 -08:00
Qi Yuhang
7ea5b42d70 [sgl-kernel][5/N]Support Expert Specialization Grouped GEMM (#12666)
Co-authored-by: Fan Yin <1106310035@qq.com>
2025-11-11 21:23:25 -08:00
zhanghaotong
5ded5e2729 [Feature] Trace: Support http/protobuf span exporter protocol (#12396)
Signed-off-by: zhanghaotong <zhanghaotong.zht@antgroup.com>
2025-11-12 05:16:43 +00:00
Mick
33d1aeb07f diffusion: refactor task type of models (#13118) 2025-11-12 12:39:34 +08:00
alisonshao
dd909a511c Fix gpt oss 4gpu b200 trace links (#12872) 2025-11-11 20:31:41 -08:00
Mick
60cb716720 [diffusion] log: improve logging while multiprocessing (#12997) 2025-11-12 12:08:37 +08:00
Trevor Morris
151e13687a Don't fuse wk+weight_proj for nextn (#12863) 2025-11-11 20:02:52 -08:00
Mick
2f9952cdbf diffusion: remove unused workflows folder (#13114) 2025-11-12 11:51:24 +08:00
MayDomine
3e7cc27318 At least tell the user that ngram verify is greedy! (#13039) 2025-11-12 11:29:33 +08:00
vipwangerxiao
8f01a12d43 Improve overlap scheduling for better TTFT (#11856)
Co-authored-by: Peng Wang <peng_wang@linux.alibaba.com>
2025-11-12 11:24:04 +08:00
yctseng0211
28b8c5792d Upgrade to ROCm 7.0 image (#13105) 2025-11-11 19:14:28 -08:00
Ziwen Zhao
7b877ab83d [Router] use call_id instead of id for matching function calls in Responses API for Harmony (#13056) 2025-11-11 17:16:47 -08:00
hlu1
0d4a418424 [Deepseek V3.2] Fix accuracy bug in the Indexer (#12583)
Signed-off-by: Hao Lu <14827759+hlu1@users.noreply.github.com>
2025-11-11 16:15:26 -08:00
Chang Su
2ca25a8aab Revert "fix: display served_model_name in /v1/models" (#13093) 2025-11-11 15:58:46 -08:00
b8zhong
cc2e36c352 overlap shared + routed expert computation in kimi linear (#12660) 2025-11-11 14:52:58 -08:00
Sai Enduri
d8f7816aba [AMD CI] Update nightly docker build CI config. (#13090) 2025-11-11 14:52:44 -08:00
Baizhou Zhang
99e25805f5 [Fix] Fix nan error for large scale ep (#12866) 2025-11-11 14:44:57 -08:00