Commit Graph

9885 Commits

Author SHA1 Message Date
Kangyan-Zhou
2bb0317e19 Remove enable_dp_attention in deepseek nightly tests (#13190) 2025-11-12 22:58:57 -08:00
fzyzcjy
86255f27b4 Revert "fallback to triton mm_persistent kernel when deepGemm fail" (#13178) 2025-11-12 22:03:35 -08:00
Hubert Lu
e4b2937017 [AMD] Add AITER Custom All-Reduce (#13102)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: HaiShaw <hixiao@gmail.com>
2025-11-12 21:53:44 -08:00
Mayyyy
7a8524b444 [Ascend] add npu synchronize (#13154)
Co-authored-by: may_feimei <meifei5@huawei.com>
2025-11-13 11:03:32 +08:00
Edmund Suen
909d0d382f docker: fix build error in Dockerfile.diffusion (#12975) 2025-11-13 09:58:16 +08:00
Tejesh Anand
e42df37df0 Remove EBNF Composer (#13163) 2025-11-12 17:55:30 -08:00
Simo Lin
6d21392b0e [router] remove worker url requirement (#13172) 2025-11-12 17:32:58 -08:00
Xiaoyu Zhang
a1cb717d0b Opt kimi_k2_thinking biased topk module (#13150) 2025-11-12 16:52:06 -08:00
Scott Lee
c94564912e Add RequestMetricsExporter utility to export request-level metrics (#10973) 2025-11-12 16:33:37 -08:00
Vladimir Serov
c4b74c1db2 [Ascend] LoRA: adding Ascend LoRA backend with using kernels from sgl_kernel_npu (#12288)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>
2025-11-12 16:20:34 -08:00
Shu Wang
7aa443903d Fix nan in global scaling factor for large scale nvfp4 EP (#13162) 2025-11-12 15:32:23 -08:00
bingps
4eda9969e8 [DeepseekV32]: use _concat_mla_absorb_q_general to replace torch.cat (#12215)
Co-authored-by: Guangda Liu <bingps@users.noreply.github.com>
2025-11-12 15:20:01 -08:00
Douglas Yang
03a7e6f4db Add job and runner failure monitor workflow for CI (#13104) 2025-11-12 14:04:50 -08:00
Xinyue Zhang
2cdde3d46d [router] Fix Flaky test_circuit_breaker_opens_and_recovers (#13164) 2025-11-12 12:07:42 -08:00
Keyang Ru
401ed0c594 [router] Add comprehensive validation to Responses API (#13127) 2025-11-12 11:19:27 -08:00
Siyuan Chen
4ef4390540 bugfix: multi-model routing for /generate api (#12979)
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
Co-authored-by: Chang Su <chang.s.su@oracle.com>
2025-11-12 11:13:27 -08:00
Lianmin Zheng
d646cf6347 [Auto Sync] Update test_deterministic.py (20251112) (#13128)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
2025-11-12 11:12:23 -08:00
Ke Bao
4edb240112 Fuse routed_scaling_factor to fused_marlin_moe (#12998) 2025-11-13 00:18:58 +08:00
Yuan Luo
706502ff6c [VLM] Support PP for Qwen2.5-VL (#13075)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn>
2025-11-12 23:18:44 +08:00
Makcum888e
c2e56dadb2 [Ascend] torch_npu.npu_mrope for MRotaryEmbedding (#10907)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>
2025-11-12 21:45:54 +08:00
Zhihao Zhang
9c1c5c6d7d [ngram] use SGLANG_NGRAM_FORCE_GREEDY_VERIFY to control verify method (#13153)
Co-authored-by: a4zhangfei <a4zhangfei@qq.com>
2025-11-12 21:07:33 +08:00
khalilzhk
2d531946db [Ascend][feature] support L1+ L2 radixcache on ascend (#12214)
Co-authored-by: khalilzhk <zhangkaikai12@huawei.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
2025-11-12 20:45:24 +08:00
sglang-bot
ebaf86d441 chore: bump SGLang version to 0.5.5.post2 (#13129)
Include the critical fix https://github.com/sgl-project/sglang/pull/12915.
2025-11-12 20:35:20 +08:00
zhanghaotong
8359f185a1 [Feature] Propagate Trace Headers into Root Span for OpenTelemetry Cross-Service Context (#10808)
Signed-off-by: zhanghaotong <zhanghaotong.zht@antgroup.com>
2025-11-12 20:28:59 +08:00
ronnie_zheng
5324f37ab3 [Ascend]adapt enable-profile-cuda-graph for NPU (#12617)
Co-authored-by: Leo920320 <1530593013@qq.com>
Co-authored-by: liupengcheng <6531487+leo920320@user.noreply.gitee.com>
2025-11-12 19:57:24 +08:00
Charlie Ruan
2864c49fbc [RPC] Fix handle_rpc_request with **recv_req.parameters (#7906) 2025-11-12 19:29:12 +08:00
sogalin
d26ec39ff2 Update aiter to v0.1.7.post1 (#13149) 2025-11-12 03:09:44 -08:00
Yikai Zhang
9c546bfd6d Fix the Wrong Return Type of Scheduler.recv_requests (#7886) 2025-11-12 18:54:56 +08:00
Rohan Potdar
c9b581644d Dump total_throughput to output-file in bench_serving.py (#9790) 2025-11-12 18:48:28 +08:00
Chang Su
e5e65e3d2a [router][grpc] Support vllm backend for grpc router (#13120) 2025-11-12 02:29:20 -08:00
SijiaYang
ffeb28ba6f fix: duplicate resize images logic of qwen-vl series models (#12458)
Signed-off-by: yangsijia.614 <yangsijia.614@bytedance.com>
2025-11-12 18:08:40 +08:00
Jimmy
b40f605fde fix(tcp-port): replace bind_server_socket to get_zmq_socket(Port conflict) (#11961)
Co-authored-by: wangchao <wcsjtu@163.com>
2025-11-12 18:05:16 +08:00
Simo Lin
3cdec20c6b [router] add minmax m2 reasoning parser (#13137) 2025-11-12 18:27:05 +09:00
Danylo Vashchilenko
d28caaf60a [router] Support complex assistant and tool messages in /chat/completions (#12860)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
2025-11-12 00:14:15 -08:00
Liangsheng Yin
ad8d24c39e Fix re-trigger actor of CI rate limit (#13136) 2025-11-12 16:10:27 +08:00
Chang Su
018123b5b0 [misc] Remove performance and router-benchmark label matching (#13135) 2025-11-12 00:04:20 -08:00
Xinyuan Tong
4983b7e7aa Fix strict level setting for Kimi K2 tool calls when not explicitly set (#13077)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2025-11-12 00:01:42 -08:00
wxsm
f825137f8a fix: Remove dupulicated kv_events initialization in scheduler (#13132) 2025-11-11 23:54:47 -08:00
Simo Lin
ae68158f2e [router] move radix tree to policy crate and addreses some code styles (#13131) 2025-11-11 23:36:04 -08:00
Ke Bao
a7cc02e36e Fix run suite sanity check (#13133) 2025-11-12 15:29:10 +08:00
Baizhou Zhang
8ece99a9dc [CI] Update job dependency and move dpsk v3.2 tests to 8-gpu suite (#12942) 2025-11-11 23:16:47 -08:00
Teng Ma
44e391b6e9 [PD] Add custom gpu id to device topo support (#12817)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2025-11-12 15:00:45 +08:00
ishandhanani
1a5c313f97 feat(engine): add rid parameter to methods in Engine class (#13095) 2025-11-11 22:37:54 -08:00
Qi Yuhang
7ea5b42d70 [sgl-kernel][5/N]Support Expert Specialization Grouped GEMM (#12666)
Co-authored-by: Fan Yin <1106310035@qq.com>
2025-11-11 21:23:25 -08:00
zhanghaotong
5ded5e2729 [Feature] Trace: Support http/protobuf span exporter protocol (#12396)
Signed-off-by: zhanghaotong <zhanghaotong.zht@antgroup.com>
2025-11-12 05:16:43 +00:00
Mick
33d1aeb07f diffusion: refactor task type of models (#13118) 2025-11-12 12:39:34 +08:00
alisonshao
dd909a511c Fix gpt oss 4gpu b200 trace links (#12872) 2025-11-11 20:31:41 -08:00
Mick
60cb716720 [diffusion] log: improve logging while multiprocessing (#12997) 2025-11-12 12:08:37 +08:00
Trevor Morris
151e13687a Don't fuse wk+weight_proj for nextn (#12863) 2025-11-11 20:02:52 -08:00
Mick
2f9952cdbf diffusion: remove unused workflows folder (#13114) 2025-11-12 11:51:24 +08:00