Commit Graph

6573 Commits

Author SHA1 Message Date
yinghui
58095cb00a Add timing metrics for requests (#12646)
Co-authored-by: Scott Lee <scottjlee@users.noreply.github.com>
2025-11-05 23:07:16 -08:00
Yuan Luo
fd3034da75 [VLM] Optimize qwen_vl preprocess_video (#12240)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: 羽癫 <yudian.zy@antgroup.com>
2025-11-06 14:55:01 +08:00
Binyao Jiang
fbbe16faab [GDN] Fuse b.sigmoid(), fused_gdn_gating and unsqueeze into one kernel: up to 0.85% e2e speedup (#12508) 2025-11-05 22:52:04 -08:00
Yi Zhang
6a1a64fa41 [BUGFIX] fix output_ids in abort (#12737) 2025-11-05 22:51:47 -08:00
Yi Zhang
e2715cf871 fix mamba prefix cache leak caused by abort (#12693) 2025-11-05 22:51:03 -08:00
Keyang Ru
74630ba310 [router][ci] Disable cache (#12752) 2025-11-05 22:49:03 -08:00
Mick
73e9a2ef5c fix: tiny fix cli (#12744) 2025-11-06 14:45:50 +08:00
Chang Su
837b08eb2e [router][grpc] Support mixin tool calls in Responses API (#12736) 2025-11-05 21:29:45 -08:00
Baizhou Zhang
bb6a21cd99 [Fix]Tiny fix in Dockerfile (#12748) 2025-11-05 21:04:33 -08:00
b8zhong
32ec68faf9 keep attention backend document up to date (#12741) 2025-11-05 20:41:20 -08:00
Baizhou Zhang
3c0a6df82d [chore] Fix triton installation for cu13 image (#12742) 2025-11-05 20:20:18 -08:00
Atream
2104d20eba Temporarily fix missing routed_scaling_factor for CompressedTensorsWNA16MoEMethod (#12738) 2025-11-06 12:01:03 +08:00
YAMY
f235498eca DeepSeek-V3.2: Add Adaptive MHA Attention Pathway for Short-Sequence Prefill (#11892) 2025-11-05 19:33:26 -08:00
alisonshao
149dc9aab1 Add nightly test multi gpu configs (#12721) 2025-11-05 19:30:19 -08:00
Baizhou Zhang
9a954982de [chore] SGLang tag management in Dockerfile (#12734) 2025-11-05 19:03:26 -08:00
gongwei-130
97be66c358 fix sgl-kernel version (#12723) 2025-11-05 19:01:03 -08:00
Keyang Ru
74243dffa2 Revert "[router] web_search_preview tool basic implementation" (#12716) 2025-11-05 18:57:39 -08:00
Keyang Ru
4ea4c48bb0 Revert "[ci] fix permission" (#12732) 2025-11-05 18:26:29 -08:00
Zaili Wang
cf5d27e30a [CPU] Upgrade default PT version to 2.9 (#12611) 2025-11-05 18:16:35 -08:00
Baizhou Zhang
9ec6031d70 [chore]Remove dockerfile from target file of bump kernel version (#12728) 2025-11-05 18:01:19 -08:00
Keyang Ru
a5affb0caf [ci] fix permission (#12729) 2025-11-05 17:43:48 -08:00
Keyang Ru
5925d3d719 fix labeler (#12718) 2025-11-05 17:34:56 -08:00
Keyang Ru
7ef1964a66 [router] add basic ci tests for gpt-oss model support (#12651)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
2025-11-05 15:45:11 -08:00
Keyang Ru
b0476a0642 [router][quick fix] Add minimal option for reasoning effort in spec (#12711) 2025-11-05 15:43:19 -08:00
Chang Su
ffba61a100 [router][grpc] Make harmony parser checks recipient first before channel (#12713) 2025-11-05 15:33:52 -08:00
Chang Su
3c219eb0e4 [misc] Change sync-labels to false (#12714) 2025-11-05 15:33:33 -08:00
Kangyan-Zhou
1ffdcdc430 Revert "Commented out b200 tests due to runner shortage (#12609)" (#12712) 2025-11-05 15:09:49 -08:00
Lianmin Zheng
c7d57d5bb3 Fix CI and style (#12658) 2025-11-05 15:08:15 -08:00
Keyang Ru
80802c4cc6 [router][ci] speed up python binding to 1.5 min (#12673) 2025-11-05 14:52:33 -08:00
Chang Su
83b104ee8c [misc] Add labeler for automatic labeling (#12710) 2025-11-05 14:44:43 -08:00
Kaixi Hou
141278048e [NVIDIA] Fix unit test of MoE and add it to nightly ci (#12709) 2025-11-05 14:33:18 -08:00
Shu Wang
82f39dc11d Add mm_fp4 trtllm backend (#12406) 2025-11-05 14:31:46 -08:00
Atream
627bac649c Support Expert Deferral Mechanism in KTransformers (#12586)
Co-authored-by: Chen Hongtao <56470055+chenht2022@users.noreply.github.com>
Co-authored-by: chenht2022 <cht22@mails.tsinghua.edu.cn>
2025-11-05 13:41:52 -08:00
wyx
3651cfbf62 [router] fix: validate HTTP status codes in health check (#12631) 2025-11-05 13:37:17 -08:00
Morpheus Guo
c8547ecddd Enable Aiter Attention for VL model (#12699)
Co-authored-by: yuechguo <yuechguo@amd.com>
2025-11-05 13:01:23 -08:00
Mick
7bc1dae095 WIP: initial multimodal-gen support (#12484)
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: yizhang2077 <1109276519@qq.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
Co-authored-by: JiLi <leege233@gmail.com>
Co-authored-by: CHEN Xi <78632976+RubiaCx@users.noreply.github.com>
Co-authored-by: laixin <xielx@shanghaitech.edu.cn>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: jzhang38 <a1286225768@gmail.com>
Co-authored-by: BrianChen1129 <yongqichcd@gmail.com>
Co-authored-by: Kevin Lin <42618777+kevin314@users.noreply.github.com>
Co-authored-by: Edenzzzz <wtan45@wisc.edu>
Co-authored-by: rlsu9 <r3su@ucsd.edu>
Co-authored-by: Jinzhe Pan <48981407+eigensystem@users.noreply.github.com>
Co-authored-by: foreverpiano <pianoqwz@qq.com>
Co-authored-by: RandNMR73 <notomatthew31@gmail.com>
Co-authored-by: PorridgeSwim <yz3883@columbia.edu>
Co-authored-by: Jiali Chen <90408393+gary-chenjl@users.noreply.github.com>
2025-11-05 12:28:52 -08:00
Chang Su
4fe53e5888 [router][grpc] Support streaming parsing with Tool Choice in chat completions API (#12677) 2025-11-05 12:12:21 -08:00
Lianmin Zheng
fb2e816e83 Fix server args for gpt oss so users can override the moe runner backend (#12696) 2025-11-05 11:36:59 -08:00
Baizhou Zhang
7c45b8b4bb [CI] Fix qwen3-vl lora nightly ci (#12708) 2025-11-05 11:00:13 -08:00
Kangyan-Zhou
ba5b68236f Commented out b200 tests due to runner shortage (#12609) 2025-11-05 10:47:43 -08:00
bigmoyan
508d2f7aa2 add Kimi k2 reasoning parser (#12702)
Signed-off-by: wangzhengtao <wangzhengtao@msh.team>
2025-11-06 00:37:54 +08:00
Yuxuan Zhang
a889c85459 [Grammar Fix] GLM-4-MOE self.first_k_dense_replace is undefined. (#12455) 2025-11-06 00:03:45 +08:00
Yuhong Guo
4d84f886e7 Refactor --debug-tensor-dump-layers to list (#12691) 2025-11-05 03:30:01 -08:00
yinghui
dc4f541823 fix trtllm_mla attention backend when disabling cuda graph. (#12687) 2025-11-05 01:35:02 -08:00
zejunchen-zejun
0648eb482d [Profiler] Add SGLANG_PROFILE_RECORD_SHAPES for recording shapes when profiling (#11641)
Signed-off-by: zejunchen-zejun <zejun.chen@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
2025-11-04 23:41:46 -08:00
yinghui
b88fab3111 fix: add seed bench_serving to cache key, remove redundant function definition. (#12680) 2025-11-04 23:39:11 -08:00
Hubert Lu
3694266051 Expand and update test coverage for AMD CI (#10044) 2025-11-04 22:15:13 -08:00
Chang Su
9f5e701879 [router][grpc] Implement tool_choice support for Responses API (#12668) 2025-11-04 21:43:42 -08:00
Glen Liu
cbf23dbbfa [Feature] add --lora-request-distribution arg to bench_serving.py and support skewed and distinct workloads (#12175) 2025-11-04 21:41:40 -08:00
Kangyan-Zhou
6dade6c3b5 Fix VLLM dependency test (#12670) 2025-11-04 20:49:58 -08:00