Commit Graph

9885 Commits

Author SHA1 Message Date
AlphaBaby
ac406d4301 fix uneven PP layer indices (#13282)
Co-authored-by: Xuchun Shang <xuchun.shang@linux.alibaba.com>
2025-11-17 19:54:23 +08:00
Simo Lin
b436113fc2 [model-gateway] Add Gateway Release Tooling (#13420) 2025-11-17 03:26:48 -08:00
Xiaoyu Zhang
8b5e2c5368 [Tiny fix] Fix bench_speculative.py run bug (#13416) 2025-11-17 18:58:19 +08:00
Simo Lin
b24235b8bb [model-gateway] update workflow names for gateway and exclude npu (#13415) 2025-11-17 02:04:07 -08:00
Liangsheng Yin
ae7698fbd5 Remove deprecated scripts (#13399) 2025-11-17 16:54:39 +08:00
Yi Zhang
a3e4fe4b41 refactor linear memory pool (#13004) 2025-11-17 16:24:29 +08:00
fzyzcjy
f3e9336dcb Support weight update for blackwell DeepGEMM (#13324) 2025-11-17 15:55:26 +08:00
Teng Ma
290fcd8971 [HiCache] add GPU id to IB dev topo for mooncake storage backend (#13112) 2025-11-17 15:48:29 +08:00
huangtingwei
1dcde53928 [HiCache] support memory_pool_host page head layout (#11644) 2025-11-17 13:45:17 +08:00
Ying Sheng
15bc1f5cd7 Update .github/MAINTAINER.md (#13398)
Co-authored-by: sglang-bot <sglangbot@gmail.com>
2025-11-16 21:32:24 -08:00
Lianmin Zheng
4d59761623 [Doc] Update CI oncall list (#13396)
Co-authored-by: sglang-bot <sglangbot@gmail.com>
Co-authored-by: Ying Sheng <sqy1415@gmail.com>
2025-11-16 20:43:12 -08:00
Liangsheng Yin
ab63f3c50b [1/N] CI refactor: introduce CI register. (#13345) 2025-11-17 12:21:20 +08:00
Zilin Zhu
147b782352 [RL] re-abort_request when model_update_lock is still locked (#13338) 2025-11-17 12:13:39 +08:00
lixiaolx
d368c7451a (1/n)support context parallel with deepseekv3.2-DSA (#12065) 2025-11-16 20:12:25 -08:00
Lianmin Zheng
7e626d12b7 Update docs (#13391)
Co-authored-by: sglang-bot <sglangbot@gmail.com>
2025-11-16 19:36:33 -08:00
Rivers
2a5773440e refactor: replace worker pool with semaphore-based concurrency in jobqueue (#13383) 2025-11-16 17:57:59 -08:00
sglang-bot
7b2fb3d47c chore: bump SGLang version to 0.5.5.post3 (#13366) 2025-11-16 17:55:38 -08:00
Baizhou Zhang
d64dd3e18e [Tiny]Fix 1-gpu nightly test bugs (#13389) 2025-11-16 15:54:17 -08:00
Baizhou Zhang
3ccd7fa669 [CI] Fix B200 CI (#13387) 2025-11-16 15:13:57 -08:00
Lifu Huang
254f62d879 Support spec decoding when LoRA is applied to target model (#12903) 2025-11-16 13:20:23 -08:00
Cheng Wan
2b8b9d8496 [CI] use cached deepep installation in gb200 CI (#13388) 2025-11-16 12:49:01 -08:00
DangKai
e970892ffa fix import qwenvl error in RL engine (#12874) 2025-11-16 12:01:16 -08:00
b8zhong
d5fa58c4dd fix nightly docker build (#13386) 2025-11-16 11:21:09 -08:00
Netanel Haber
9f011f617f fix generative_models.md table - remove newlines (#13385) 2025-11-16 10:33:38 -08:00
ybyang
3b18fd4cf5 [router] bindings for go (#13384)
Signed-off-by: ybyang <ybyang7@iflytek.com>
2025-11-16 10:12:10 -08:00
Liangsheng Yin
1869f25ca5 Tiny deprecate the --range-begin in run_suite.py (#13381) 2025-11-16 23:15:36 +08:00
Liangsheng Yin
e019f233f9 Remove unused code / testcases in lang (#13335) 2025-11-16 22:48:35 +08:00
Xinyuan Tong
b1c688fba2 refactor: cleanup vision attention related codes (#13228)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: alisonshao <54658187+alisonshao@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2025-11-16 22:44:40 +08:00
Zilin Zhu
8e3663d4e8 Add default enable_memory_saver to HybridLinearKVPool (#13371) 2025-11-16 22:35:13 +08:00
Zilin Zhu
9edb0e0da3 Add SGLANG_ENABLE_REQ_POOL_LEAK_STRICT_CHECK to bypass mem leak check (#13339) 2025-11-16 22:02:22 +08:00
Xiaoyu Zhang
95f43669b5 [2 / 2] apply sgl-kernel weak_ref_tensor (#12978)
Co-authored-by: Yuan Luo <yuan.luo@hotmail.com>
2025-11-16 21:24:35 +08:00
Xiaoyu Zhang
50691d7b49 [opt kimi k2 2/n] apply kimi k2 thinking moe_fused_gate (#13332) 2025-11-16 21:20:08 +08:00
Yuhao Yang
6afe396399 diffusion: support fa4 in fa backend for blackwell (#13263)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-11-16 21:02:45 +08:00
Mick
191f5c7795 diffusion: correct check-changes for multimodal_gen (#13375) 2025-11-16 20:32:33 +08:00
Adarsh Shirawalmath
d724670873 [diffusion] refactor and added tests for Flux, T2V, TI2V, I2V(#13344) 2025-11-16 19:58:04 +08:00
Simo Lin
efc5d8f5ed [model-gateway] fix SDist step readme path (#13373) 2025-11-16 03:16:52 -08:00
Simo Lin
ef32a252d0 [model-gateway] fix model gateway pypi release workflow path (#13372) 2025-11-16 02:33:49 -08:00
Zilin Zhu
9509c4ccd3 [RL] enable offloading hybrid linear attn model (#13336) 2025-11-16 17:12:52 +08:00
Kangyan-Zhou
4e19c1d541 Add missing models (#13369) 2025-11-16 00:04:47 -08:00
Qiaolin Yu
78a4b446c6 Fix dpsk-r1-fp4 tp8 by reverting two commits (#13162 and #13341) (#13348)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2025-11-15 21:31:36 -08:00
DarkSharpness
f969664172 [Performance] Move the contiguous to torch compile region (#13199) 2025-11-15 20:49:52 -08:00
b8zhong
f35f7f1245 [Piecewise CUDA Graph] Support W4A8 (#13179) 2025-11-16 11:53:50 +08:00
ash-sigh
12c789ebd1 [Ascend]support xgrammar backend for ascend npu (#12310)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2025-11-15 19:39:24 -08:00
kyleliang-nv
597d416070 [feature] Add layerwise NVTX support (#11870) 2025-11-15 19:20:56 -08:00
sglang-bot
1ca205f6da chore: bump sgl-kernel version to 0.3.17.post1 (#13358) 2025-11-15 19:11:41 -08:00
b8zhong
24a25ffa20 [Piecewise CUDA Graph] Support ModelOpt FP4 (#13101) 2025-11-16 11:03:19 +08:00
Lianmin Zheng
db7299aa30 [Fix] Register custom ops only if they exist (#13321) 2025-11-15 17:48:17 -08:00
Cheng Wan
13366843a4 [CI] check unit-test-backend-8-gpu-h20 in workflow (#13355) 2025-11-15 16:44:10 -08:00
Kangyan-Zhou
be353ffd13 Add missing models (#13351) 2025-11-15 13:51:20 -08:00
iLeGend
20e59f9510 Add FP32 dtype support for RoPE - Part1 (#13181) 2025-11-15 11:37:18 -08:00