Commit Graph
6126 Commits
Author SHA1 Message Date
RangerCD a480ca7ead fix: zmq_to_tokenizer encoder transfer when host listens to 0.0.0.0 (#17929) 2026-02-02 13:27:27 +08:00
siyu f1824a957b [EPD][refactor]: introduce BaseMMReceiver for gRPC transport integration (#17921) 2026-02-02 11:37:32 +08:00
Cheng Wan ab8b99eb23 Refine logprob logic for request handling (#17986) 2026-02-01 19:11:52 -08:00
8ed35df204 Add bootstrap_room validation to detect metadata corruption in PD disaggregation (#17430)
Co-authored-by: 继优 <jiyou.ljy@alibaba-inc.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-02-02 10:43:50 +08:00
Mick c84cd4b5ff [diffusion] fix: fix missing component names for VAELoader (#18069) 2026-02-02 09:48:17 +08:00
Mick 977096ae03 [diffusion] cli: introduce generic attention backend configuration in ServerArgs (#18036) 2026-02-02 09:47:40 +08:00
Yuhao Yang d11ccc0a0a fix: avoid double reduce in VLM dp attention (#17991) 2026-02-02 09:44:32 +08:00
Yuan Luoandluoyuan.luo 9227d4f748 [Fix] Remove no use code in MiMo-V2-Flash (#18051)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-02-01 15:34:09 -08:00
Glen Liu 99dad105fd [TestFix] rewrite LoRA overlap loading tests (#18047) 2026-02-01 14:52:08 -08:00
Koushik Duttaandroot 993ec178ef [BUGFIX]: using language-only should not reserve space for the vision encoder (#18011)
Co-authored-by: root <root@ubuntu-nvidia.localdomain>
2026-02-01 14:49:00 -08:00
Lianmin Zhenggithub-actions[bot] <github-actions[bot]@users.noreply.github.com>Jiayi Yuangemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
1acae30806 [Auto Sync] Update test_deterministic.py (20260131) (#18034)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Jiayi Yuan <34369239+jy-yuan@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-01 07:40:07 -08:00
Lianmin Zhengandgithub-actions[bot] <github-actions[bot]@users.noreply.github.com> fb609669ca [Auto Sync] Update elementwise.py (20260131) (#18033)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-02-01 07:39:37 -08:00
Yuan Luoandluoyuan.luo 4ea4f2a20c [VLM] Optimize get_rope_index for GLM4v (#17420)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-02-01 18:59:15 +08:00
jiashaokun-1 0fe282543f [NPU] support the Enable return routed experts (#17025) 2026-02-01 18:31:39 +08:00
27bec34203 [NPU] disaggregation_decode_enable_fake_auto parameter adaptation (#17811)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
2026-02-01 18:14:26 +08:00
Ke Bao d9050b4a9c Reset evict swa status when retract (#18059) 2026-02-01 17:17:37 +08:00
47592a23c7 [CI] Fix AMD CI by inlining dummy_grok config (#18044)
Co-authored-by: root <root@mi300x8-005.atl1.do.cpe.ice.amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-01 00:20:57 -08:00
855dd0546c feat: Add Ling Flash v2.0 support for Eagle3 (#15119)
Co-authored-by: chenyefei.cyf <chenyefei.cyf@U-9V5T77LW-2356.local>
Co-authored-by: GeLee-Q <865038696@qq.com>
Co-authored-by: Gao016 <yngao016@163.com>
Co-authored-by: Shenggui Li <somerlee.9@gmail.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
2026-01-31 23:57:45 -08:00
lukec 3ca29dffc7 support qwen3-next eagle3 (#14607) 2026-01-31 23:45:23 -08:00
Praneth ParuchuriandYuhao Yang 9bb1260558 [Feature] Support file:// URL format for multimodal inputs (#14490)
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
2026-01-31 23:44:44 -08:00
Roger Young 486c7de39f Optimizing all_reduce in RMSNormTP in minimax_m2 (#16483) 2026-01-31 23:39:37 -08:00
Kangyan-Zhou 9c168fcac7 Fix Diffusion Request Validation to allow missing input artifacts if the input only contains text (#16610) 2026-01-31 23:38:40 -08:00
tc-mbandXinyuan Tong 4d28cda007 [model] Support MiniCPM-V 4.5 (#9610)
Signed-off-by: tc-mb <caitianchi@modelbest.cn>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-01-31 23:37:36 -08:00
linhaifeng 2c036f1eb1 [Bugfix] fix the display error (inconsistent context) (#17699)
Signed-off-by: linhaifeng <1371675203@qq.com>
2026-01-31 23:35:11 -08:00
Ke Bao d396650bd2 Fix swa kv cache memory allocation (#18039) 2026-02-01 14:26:51 +08:00
Minglei Zhu 38d275a9fd [BugFix] fix gpt-oss accuracy issue when enabling piecewise cuda graph (#18013) 2026-02-01 14:26:26 +08:00
Yinghai Lu 11892599f1 [metric] Optional extra metric labels (#18049) 2026-02-01 14:25:59 +08:00
Baizhou Zhang c7d53fa26a Set torch url index in pyproject.toml (#16802) 2026-02-01 13:23:52 +08:00
khalilzhk 429ef988bc [BugFix] Fix draft model specified config file (#17815) 2026-01-31 20:45:36 -08:00
Kaixi 2b2515423a Skipped warning on sm100 (#18000) 2026-01-31 20:21:03 -08:00
Kangyan-Zhou d443d2d2ae Improve error output in tnightly tets (#18053) 2026-01-31 19:26:19 -08:00
0f2df9370a feat: validate ib devices in server args (#17598)
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
2026-01-31 17:50:42 -08:00
b8zhongandBrayden Zhong 398d13a189 [Perf] Add Flashinfer DeepGEMM SM90 for SwapAB Optimization (#15514)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
2026-02-01 08:56:23 +08:00
Chongchong Tian 9951a1ae07 Fix: Remove duplicate assignment for use_w4afp8 (#17858) 2026-01-31 16:04:08 -08:00
Yi Zhong c2ab3713e9 [Performance] Optimize Mllama LayerNorm -> Upd (#9725) 2026-01-31 16:02:57 -08:00
Hao Jinandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> 6aaea09b3d Update python/sglang/README.md (#18045)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-31 13:20:41 -08:00
Zaili Wang 97593c9f41 [CPU] toml file update (#17861) 2026-01-31 13:16:06 -08:00
Mohammad Miadh Angkad 9ac4dcada4 [Tiny] Fix grammar in shared experts fusion log messages (#18043) 2026-01-31 13:14:25 -08:00
b8zhong ef134d407d [Fix] Revert back to using CUTLASS mm_fp4 backend (#17369) 2026-01-31 23:01:29 +08:00
Mick 1a006c2a0d [diffusion] refactor: split component_loader into component-wise files (#17820) 2026-01-31 20:22:31 +08:00
Lianmin Zhenggithub-actions[bot] <github-actions[bot]@users.noreply.github.com>Byron Hsu
7412ceb4eb [Auto Sync] Update linear.py to assert shapes (20260130) (#17966)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
2026-01-31 01:01:55 -08:00
Lianmin Zhenggithub-actions[bot] <github-actions[bot]@users.noreply.github.com>Archit Patke
0e184609d3 Add launch_command assignment in crash dump (#17967)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Archit Patke <apatke@x.ai>
2026-01-31 01:00:40 -08:00
b8zhong 22498e10c0 [Fix] Triton TP MoE Dpsk V3/Qwen3 Coder with SwapAB (#17965) 2026-01-31 15:56:26 +08:00
Zheng Wengangandsiyu a4df95c15f [EPD][Perf] parallelize ZMQ send for encode server (#16487)
Co-authored-by: siyu <liusy58@linux.alibaba.com>
2026-01-31 14:30:11 +08:00
jeff 04efd03dbf Fix OOM in DeepSeek weight loading by deferring dict(weights) materialization (#17744) 2026-01-31 13:59:00 +08:00
Hudson Xing c72bf50706 add reasoning_tokens usage test for tool call (#18022) 2026-01-30 21:09:23 -08:00
Mohammad Miadh Angkad d0d9cecd1b Fix cuBLAS >=12.9 detection for cu12/cu13 package naming (#17766) 2026-01-31 12:01:52 +08:00
Xiaoyu Zhang 22aad4e2c4 [Diffusion] Fix FLUX.1-schnell time embedding argument mismatch (#17988) 2026-01-31 11:47:27 +08:00
Bi Xue 5d00150e99 [sglang] fix mm token padded value overlap with text token id (#17781) 2026-01-30 17:09:13 -08:00
e86476acfc [NPU] support llama-3.2-11B-vision-instruct mode for NPU (#17492)
Co-authored-by: McZyWu <zhuoyun.wu.23@ucl.ac.uk>
Co-authored-by: chenyang08056032 <chenyang08056032@163.com>
Co-authored-by: Hexq0210 <893781835@qq.com>
2026-01-31 08:49:38 +08:00