Commit Graph
10693 Commits
Author SHA1 Message Date
Liangsheng Yin 85b8c5c4cd Set max parallel for 1-gpu runner (#13215) 2025-11-14 01:36:43 +08:00
neo c8b7516fe8 Fix accept rate in speculative decoding metrics (#13212) 2025-11-14 00:54:50 +08:00
67e9d287ee [Quantization] Support Quark Dense + MoE FP8 & FP8 PTPC (#10485)
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
2025-11-13 08:16:00 -08:00
Sam e7e89349c9 Enable Flashinfer TRTLLM-GEN-MoE FP8 blockwise kernel for Qwen3-Next on Blackwell (#12543) 2025-11-13 19:44:44 +08:00
aead0ef5e5 [FEAT][ROCM] enable fused shared expert for Rocm (#12201)
Co-authored-by: ZLkanyo009 <4071250045@qq.com>
Co-authored-by: HAI <hixiao@gmail.com>
2025-11-13 00:41:40 -08:00
Shu Wang 6664083522 Replace [silu_and_mul_]scaled_fp4_group_quant by Flashinfer equivalent (#12376) 2025-11-13 00:26:00 -08:00
Edmund Suen c2d69e8b56 fix: display served_model_name in /v1/models (#13155) 2025-11-13 17:07:55 +09:00
Simo Lin 4c1e909a80 [router] minmax-m2 xml tool parser (#13148) 2025-11-12 23:24:54 -08:00
Kangyan-Zhou 2bb0317e19 Remove enable_dp_attention in deepseek nightly tests (#13190) 2025-11-12 22:58:57 -08:00
fzyzcjy 86255f27b4 Revert "fallback to triton mm_persistent kernel when deepGemm fail" (#13178) 2025-11-12 22:03:35 -08:00
e4b2937017 [AMD] Add AITER Custom All-Reduce (#13102)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: HaiShaw <hixiao@gmail.com>
2025-11-12 21:53:44 -08:00
Mayyyyandmay_feimei 7a8524b444 [Ascend] add npu synchronize (#13154)
Co-authored-by: may_feimei <meifei5@huawei.com>
2025-11-13 11:03:32 +08:00
Edmund Suen 909d0d382f docker: fix build error in Dockerfile.diffusion (#12975) 2025-11-13 09:58:16 +08:00
Tejesh Anand e42df37df0 Remove EBNF Composer (#13163) 2025-11-12 17:55:30 -08:00
Simo Lin 6d21392b0e [router] remove worker url requirement (#13172) 2025-11-12 17:32:58 -08:00
Xiaoyu Zhang a1cb717d0b Opt kimi_k2_thinking biased topk module (#13150) 2025-11-12 16:52:06 -08:00
Scott Lee c94564912e Add RequestMetricsExporter utility to export request-level metrics (#10973) 2025-11-12 16:33:37 -08:00
c4b74c1db2 [Ascend] LoRA: adding Ascend LoRA backend with using kernels from sgl_kernel_npu (#12288)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>
2025-11-12 16:20:34 -08:00
Shu Wang 7aa443903d Fix nan in global scaling factor for large scale nvfp4 EP (#13162) 2025-11-12 15:32:23 -08:00
bingpsandGuangda Liu 4eda9969e8 [DeepseekV32]: use _concat_mla_absorb_q_general to replace torch.cat (#12215)
Co-authored-by: Guangda Liu <bingps@users.noreply.github.com>
2025-11-12 15:20:01 -08:00
Douglas Yang 03a7e6f4db Add job and runner failure monitor workflow for CI (#13104) 2025-11-12 14:04:50 -08:00
Xinyue Zhang 2cdde3d46d [router] Fix Flaky test_circuit_breaker_opens_and_recovers (#13164) 2025-11-12 12:07:42 -08:00
Keyang Ru 401ed0c594 [router] Add comprehensive validation to Responses API (#13127) 2025-11-12 11:19:27 -08:00
4ef4390540 bugfix: multi-model routing for /generate api (#12979)
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
Co-authored-by: Chang Su <chang.s.su@oracle.com>
2025-11-12 11:13:27 -08:00
Lianmin ZhengandStefan He d646cf6347 [Auto Sync] Update test_deterministic.py (20251112) (#13128)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
2025-11-12 11:12:23 -08:00
Ke Bao 4edb240112 Fuse routed_scaling_factor to fused_marlin_moe (#12998) 2025-11-13 00:18:58 +08:00
706502ff6c [VLM] Support PP for Qwen2.5-VL (#13075)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn>
2025-11-12 23:18:44 +08:00
c2e56dadb2 [Ascend] torch_npu.npu_mrope for MRotaryEmbedding (#10907)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>
2025-11-12 21:45:54 +08:00
Zhihao Zhanganda4zhangfei 9c1c5c6d7d [ngram] use SGLANG_NGRAM_FORCE_GREEDY_VERIFY to control verify method (#13153)
Co-authored-by: a4zhangfei <a4zhangfei@qq.com>
2025-11-12 21:07:33 +08:00
2d531946db [Ascend][feature] support L1+ L2 radixcache on ascend (#12214)
Co-authored-by: khalilzhk <zhangkaikai12@huawei.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
2025-11-12 20:45:24 +08:00
sglang-bot ebaf86d441 chore: bump SGLang version to 0.5.5.post2 (#13129)
Include the critical fix https://github.com/sgl-project/sglang/pull/12915.
2025-11-12 20:35:20 +08:00
zhanghaotong 8359f185a1 [Feature] Propagate Trace Headers into Root Span for OpenTelemetry Cross-Service Context (#10808)
Signed-off-by: zhanghaotong <zhanghaotong.zht@antgroup.com>
2025-11-12 20:28:59 +08:00
5324f37ab3 [Ascend]adapt enable-profile-cuda-graph for NPU (#12617)
Co-authored-by: Leo920320 <1530593013@qq.com>
Co-authored-by: liupengcheng <6531487+leo920320@user.noreply.gitee.com>
2025-11-12 19:57:24 +08:00
Charlie Ruan 2864c49fbc [RPC] Fix handle_rpc_request with **recv_req.parameters (#7906) 2025-11-12 19:29:12 +08:00
sogalin d26ec39ff2 Update aiter to v0.1.7.post1 (#13149) 2025-11-12 03:09:44 -08:00
Yikai Zhang 9c546bfd6d Fix the Wrong Return Type of Scheduler.recv_requests (#7886) 2025-11-12 18:54:56 +08:00
Rohan Potdar c9b581644d Dump total_throughput to output-file in bench_serving.py (#9790) 2025-11-12 18:48:28 +08:00
Chang Su e5e65e3d2a [router][grpc] Support vllm backend for grpc router (#13120) 2025-11-12 02:29:20 -08:00
SijiaYang ffeb28ba6f fix: duplicate resize images logic of qwen-vl series models (#12458)
Signed-off-by: yangsijia.614 <yangsijia.614@bytedance.com>
2025-11-12 18:08:40 +08:00
Jimmyandwangchao b40f605fde fix(tcp-port): replace bind_server_socket to get_zmq_socket(Port conflict) (#11961)
Co-authored-by: wangchao <wcsjtu@163.com>
2025-11-12 18:05:16 +08:00
Simo Lin 3cdec20c6b [router] add minmax m2 reasoning parser (#13137) 2025-11-12 18:27:05 +09:00
d28caaf60a [router] Support complex assistant and tool messages in /chat/completions (#12860)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
2025-11-12 00:14:15 -08:00
Liangsheng Yin ad8d24c39e Fix re-trigger actor of CI rate limit (#13136) 2025-11-12 16:10:27 +08:00
Chang Su 018123b5b0 [misc] Remove performance and router-benchmark label matching (#13135) 2025-11-12 00:04:20 -08:00
Xinyuan Tong 4983b7e7aa Fix strict level setting for Kimi K2 tool calls when not explicitly set (#13077)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2025-11-12 00:01:42 -08:00
wxsm f825137f8a fix: Remove dupulicated kv_events initialization in scheduler (#13132) 2025-11-11 23:54:47 -08:00
Simo Lin ae68158f2e [router] move radix tree to policy crate and addreses some code styles (#13131) 2025-11-11 23:36:04 -08:00
Ke Bao a7cc02e36e Fix run suite sanity check (#13133) 2025-11-12 15:29:10 +08:00
Baizhou Zhang 8ece99a9dc [CI] Update job dependency and move dpsk v3.2 tests to 8-gpu suite (#12942) 2025-11-11 23:16:47 -08:00
Teng MaandShangming Cai 44e391b6e9 [PD] Add custom gpu id to device topo support (#12817)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2025-11-12 15:00:45 +08:00