Kangyan-Zhou
|
2bb0317e19
|
Remove enable_dp_attention in deepseek nightly tests (#13190)
|
2025-11-12 22:58:57 -08:00 |
|
fzyzcjy
|
86255f27b4
|
Revert "fallback to triton mm_persistent kernel when deepGemm fail" (#13178)
|
2025-11-12 22:03:35 -08:00 |
|
Hubert Lu
|
e4b2937017
|
[AMD] Add AITER Custom All-Reduce (#13102)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2025-11-12 21:53:44 -08:00 |
|
Mayyyy
|
7a8524b444
|
[Ascend] add npu synchronize (#13154)
Co-authored-by: may_feimei <meifei5@huawei.com>
|
2025-11-13 11:03:32 +08:00 |
|
Edmund Suen
|
909d0d382f
|
docker: fix build error in Dockerfile.diffusion (#12975)
|
2025-11-13 09:58:16 +08:00 |
|
Tejesh Anand
|
e42df37df0
|
Remove EBNF Composer (#13163)
|
2025-11-12 17:55:30 -08:00 |
|
Simo Lin
|
6d21392b0e
|
[router] remove worker url requirement (#13172)
|
2025-11-12 17:32:58 -08:00 |
|
Xiaoyu Zhang
|
a1cb717d0b
|
Opt kimi_k2_thinking biased topk module (#13150)
|
2025-11-12 16:52:06 -08:00 |
|
Scott Lee
|
c94564912e
|
Add RequestMetricsExporter utility to export request-level metrics (#10973)
|
2025-11-12 16:33:37 -08:00 |
|
Vladimir Serov
|
c4b74c1db2
|
[Ascend] LoRA: adding Ascend LoRA backend with using kernels from sgl_kernel_npu (#12288)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>
|
2025-11-12 16:20:34 -08:00 |
|
Shu Wang
|
7aa443903d
|
Fix nan in global scaling factor for large scale nvfp4 EP (#13162)
|
2025-11-12 15:32:23 -08:00 |
|
bingps
|
4eda9969e8
|
[DeepseekV32]: use _concat_mla_absorb_q_general to replace torch.cat (#12215)
Co-authored-by: Guangda Liu <bingps@users.noreply.github.com>
|
2025-11-12 15:20:01 -08:00 |
|
Douglas Yang
|
03a7e6f4db
|
Add job and runner failure monitor workflow for CI (#13104)
|
2025-11-12 14:04:50 -08:00 |
|
Xinyue Zhang
|
2cdde3d46d
|
[router] Fix Flaky test_circuit_breaker_opens_and_recovers (#13164)
|
2025-11-12 12:07:42 -08:00 |
|
Keyang Ru
|
401ed0c594
|
[router] Add comprehensive validation to Responses API (#13127)
|
2025-11-12 11:19:27 -08:00 |
|
Siyuan Chen
|
4ef4390540
|
bugfix: multi-model routing for /generate api (#12979)
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
Co-authored-by: Chang Su <chang.s.su@oracle.com>
|
2025-11-12 11:13:27 -08:00 |
|
Lianmin Zheng
|
d646cf6347
|
[Auto Sync] Update test_deterministic.py (20251112) (#13128)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
|
2025-11-12 11:12:23 -08:00 |
|
Ke Bao
|
4edb240112
|
Fuse routed_scaling_factor to fused_marlin_moe (#12998)
|
2025-11-13 00:18:58 +08:00 |
|
Yuan Luo
|
706502ff6c
|
[VLM] Support PP for Qwen2.5-VL (#13075)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn>
|
2025-11-12 23:18:44 +08:00 |
|
Makcum888e
|
c2e56dadb2
|
[Ascend] torch_npu.npu_mrope for MRotaryEmbedding (#10907)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: ssshinigami <44640852+ssshinigami@users.noreply.github.com>
|
2025-11-12 21:45:54 +08:00 |
|
Zhihao Zhang
|
9c1c5c6d7d
|
[ngram] use SGLANG_NGRAM_FORCE_GREEDY_VERIFY to control verify method (#13153)
Co-authored-by: a4zhangfei <a4zhangfei@qq.com>
|
2025-11-12 21:07:33 +08:00 |
|
khalilzhk
|
2d531946db
|
[Ascend][feature] support L1+ L2 radixcache on ascend (#12214)
Co-authored-by: khalilzhk <zhangkaikai12@huawei.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
|
2025-11-12 20:45:24 +08:00 |
|
sglang-bot
|
ebaf86d441
|
chore: bump SGLang version to 0.5.5.post2 (#13129)
Include the critical fix https://github.com/sgl-project/sglang/pull/12915.
|
2025-11-12 20:35:20 +08:00 |
|
zhanghaotong
|
8359f185a1
|
[Feature] Propagate Trace Headers into Root Span for OpenTelemetry Cross-Service Context (#10808)
Signed-off-by: zhanghaotong <zhanghaotong.zht@antgroup.com>
|
2025-11-12 20:28:59 +08:00 |
|
ronnie_zheng
|
5324f37ab3
|
[Ascend]adapt enable-profile-cuda-graph for NPU (#12617)
Co-authored-by: Leo920320 <1530593013@qq.com>
Co-authored-by: liupengcheng <6531487+leo920320@user.noreply.gitee.com>
|
2025-11-12 19:57:24 +08:00 |
|
Charlie Ruan
|
2864c49fbc
|
[RPC] Fix handle_rpc_request with **recv_req.parameters (#7906)
|
2025-11-12 19:29:12 +08:00 |
|
sogalin
|
d26ec39ff2
|
Update aiter to v0.1.7.post1 (#13149)
|
2025-11-12 03:09:44 -08:00 |
|
Yikai Zhang
|
9c546bfd6d
|
Fix the Wrong Return Type of Scheduler.recv_requests (#7886)
|
2025-11-12 18:54:56 +08:00 |
|
Rohan Potdar
|
c9b581644d
|
Dump total_throughput to output-file in bench_serving.py (#9790)
|
2025-11-12 18:48:28 +08:00 |
|
Chang Su
|
e5e65e3d2a
|
[router][grpc] Support vllm backend for grpc router (#13120)
|
2025-11-12 02:29:20 -08:00 |
|
SijiaYang
|
ffeb28ba6f
|
fix: duplicate resize images logic of qwen-vl series models (#12458)
Signed-off-by: yangsijia.614 <yangsijia.614@bytedance.com>
|
2025-11-12 18:08:40 +08:00 |
|
Jimmy
|
b40f605fde
|
fix(tcp-port): replace bind_server_socket to get_zmq_socket(Port conflict) (#11961)
Co-authored-by: wangchao <wcsjtu@163.com>
|
2025-11-12 18:05:16 +08:00 |
|
Simo Lin
|
3cdec20c6b
|
[router] add minmax m2 reasoning parser (#13137)
|
2025-11-12 18:27:05 +09:00 |
|
Danylo Vashchilenko
|
d28caaf60a
|
[router] Support complex assistant and tool messages in /chat/completions (#12860)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
|
2025-11-12 00:14:15 -08:00 |
|
Liangsheng Yin
|
ad8d24c39e
|
Fix re-trigger actor of CI rate limit (#13136)
|
2025-11-12 16:10:27 +08:00 |
|
Chang Su
|
018123b5b0
|
[misc] Remove performance and router-benchmark label matching (#13135)
|
2025-11-12 00:04:20 -08:00 |
|
Xinyuan Tong
|
4983b7e7aa
|
Fix strict level setting for Kimi K2 tool calls when not explicitly set (#13077)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-11-12 00:01:42 -08:00 |
|
wxsm
|
f825137f8a
|
fix: Remove dupulicated kv_events initialization in scheduler (#13132)
|
2025-11-11 23:54:47 -08:00 |
|
Simo Lin
|
ae68158f2e
|
[router] move radix tree to policy crate and addreses some code styles (#13131)
|
2025-11-11 23:36:04 -08:00 |
|
Ke Bao
|
a7cc02e36e
|
Fix run suite sanity check (#13133)
|
2025-11-12 15:29:10 +08:00 |
|
Baizhou Zhang
|
8ece99a9dc
|
[CI] Update job dependency and move dpsk v3.2 tests to 8-gpu suite (#12942)
|
2025-11-11 23:16:47 -08:00 |
|
Teng Ma
|
44e391b6e9
|
[PD] Add custom gpu id to device topo support (#12817)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2025-11-12 15:00:45 +08:00 |
|
ishandhanani
|
1a5c313f97
|
feat(engine): add rid parameter to methods in Engine class (#13095)
|
2025-11-11 22:37:54 -08:00 |
|
Qi Yuhang
|
7ea5b42d70
|
[sgl-kernel][5/N]Support Expert Specialization Grouped GEMM (#12666)
Co-authored-by: Fan Yin <1106310035@qq.com>
|
2025-11-11 21:23:25 -08:00 |
|
zhanghaotong
|
5ded5e2729
|
[Feature] Trace: Support http/protobuf span exporter protocol (#12396)
Signed-off-by: zhanghaotong <zhanghaotong.zht@antgroup.com>
|
2025-11-12 05:16:43 +00:00 |
|
Mick
|
33d1aeb07f
|
diffusion: refactor task type of models (#13118)
|
2025-11-12 12:39:34 +08:00 |
|
alisonshao
|
dd909a511c
|
Fix gpt oss 4gpu b200 trace links (#12872)
|
2025-11-11 20:31:41 -08:00 |
|
Mick
|
60cb716720
|
[diffusion] log: improve logging while multiprocessing (#12997)
|
2025-11-12 12:08:37 +08:00 |
|
Trevor Morris
|
151e13687a
|
Don't fuse wk+weight_proj for nextn (#12863)
|
2025-11-11 20:02:52 -08:00 |
|
Mick
|
2f9952cdbf
|
diffusion: remove unused workflows folder (#13114)
|
2025-11-12 11:51:24 +08:00 |
|