Commit Graph

11116 Commits

Author SHA1 Message Date
Shangming Cai
93452a8252 [PD] Support decode pp for PD disaggregation (#14265)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2025-12-03 14:35:29 +08:00
b8zhong
65c8568c4a sync attention, deepseek doc (#14335)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
2025-12-02 21:19:40 -08:00
Baizhou Zhang
4bcc5879af [Doc] Fix DeepSeek V32 Doc (#14336) 2025-12-02 21:06:55 -08:00
Simo Lin
d5ea8c7176 [model-gateway] multimodality initialization (#13350) 2025-12-02 19:41:40 -08:00
Vikram
42271376d1 [bug fix] use npu phy id in container env (#14266)
Co-authored-by: jinke15 <jinke15@jd.com>
2025-12-03 11:33:43 +08:00
Mick
84e0abb7b1 Update CODEOWNERS for multimodal (#14329) 2025-12-03 11:04:40 +08:00
Johnsonms
043f13171f [Performance] Optimize NSA Indexer K/S Buffer Access with Fused Triton Kernels (#13812)
Co-authored-by: Johnsonms <johnson@together.ai>
2025-12-02 18:53:06 -08:00
Dongjie Zou
f764c6910d [diffusion] feat: support distilled vae generic (#14195)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-12-03 10:27:31 +08:00
Baizhou Zhang
922054079c [Doc] Update DeepSeek-V3.2 document (#14321) 2025-12-02 18:19:39 -08:00
Even Zhou
7d1a130cde Refactor custom allreduce logics (#13710)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2025-12-02 17:20:05 -08:00
sglang-bot
7ae368efde chore: bump SGLang version to 0.5.6 (#14316)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2025-12-02 17:17:13 -08:00
sunxxuns
c4e20caded ci: Add zyzshishui to CI permissions (#14324)
Co-authored-by: root <root@mi300x8-005.atl1.do.cpe.ice.amd.com>
2025-12-02 17:00:31 -08:00
Simo Lin
5dad1ff1b6 [model-gateway] add workflow for external model providers (#14323) 2025-12-02 16:47:06 -08:00
Lianmin Zheng
ca52ed425f Clean up imports and move files (#14317) 2025-12-02 16:31:54 -08:00
Douglas Yang
5c03aa3e9d Adding section for scheduled PR test runs on main (#14309) 2025-12-02 14:52:44 -08:00
Douglas Yang
253be18e52 Fix nonetype error for ci failure monitor (#14319) 2025-12-02 14:24:25 -08:00
alisonshao
084b06e79d Add /rerun-stage slash command to rerun specific PR test stages (#14262) 2025-12-02 14:23:38 -08:00
Lianmin Zheng
fc6fb550ec [Minor] update docs on CI (#14315) 2025-12-02 12:06:54 -08:00
Eva20150932-atlascloud
7c38eca1e4 feat: DeepSeek new v3.2 encoding (#14249)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2025-12-02 11:41:05 -08:00
Quanfeng Li
427b08e24d Init TBO with dp_padded batch (#11423)
Co-authored-by: Cheng Wan <wan4ch@gmail.com>
Co-authored-by: Yuhao Yao <37280700+yuhyao@users.noreply.github.com>
2025-12-02 10:34:26 -08:00
alisonshao
0141ca370f Revert PR #14044: Restore separate memory pool for piecewise CUDA graph (#14278) 2025-12-02 09:53:16 -08:00
alisonshao
25a6be4930 Fix duplicate download log messages in multi-process environment (#14299) 2025-12-02 09:33:18 -08:00
Antonin Vidon
df1f31241b [sgl-kernel] fix runtime error while preloading CUDA runtime (#13089) 2025-12-02 23:03:51 +08:00
Xiaoyu Zhang
c5947ecd85 Opt moe align block size kernel (#14133) 2025-12-02 19:13:55 +08:00
Mick
9530b76630 [diffusion] refactor: simplify DmdDenoisingStage (#14269) 2025-12-02 18:59:40 +08:00
Jinyan Chen
3067b3f050 [diffusion] chore: improve model info registration and searching strategy (#14281)
Co-authored-by: Jinyan Chen <jinyanc@nvidia.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-12-02 18:28:59 +08:00
alisonshao
e0ec42c710 [CI] Fix 4-GPU test timeout by using 3 partitions (#14287) 2025-12-02 16:47:26 +08:00
Mick
9c9d7091cb Update CODEOWNERS for multimodal_gen (#14286) 2025-12-02 16:15:19 +08:00
Simo Lin
51a86ce69e [model-gateway] change rust package name to sgl-model-gateway instead (#14283) 2025-12-02 00:06:54 -08:00
Lianmin Zheng
64092c8b55 [Auto Sync] Rename is_hybrid to is_hybrid_swa (#14252)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: Hanming Lu <hanming@x.ai>
2025-12-01 23:24:24 -08:00
sglang-bot
63b9300f00 chore: bump sgl-kernel version to 0.3.18.post2 (#14244) 2025-12-01 23:14:12 -08:00
Yuan Luo
21ec99beff [VLM][Doc] Document for VLM DP Encoder (#14279)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-12-02 15:08:37 +08:00
Simo Lin
383689e3ad [model-gateway] fix version output (#14276) 2025-12-01 22:51:57 -08:00
b8zhong
236a7c2370 fix trtllm mla spec (#13738)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
2025-12-01 22:16:25 -08:00
Roger Young
3dabd609fb Optimize topk sigmoid in minimax_m2 (#14047)
Co-authored-by: xuebi <xuebi@minimaxi.com>
2025-12-02 14:07:12 +08:00
Simo Lin
73df525382 [model-gateway] include smg version command in py binding (#14274) 2025-12-01 21:21:02 -08:00
b8zhong
e6420100ee sync attention doc and ep doc to doctree (#14257)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
2025-12-01 21:15:22 -08:00
Kevin Li
c9e2090101 fix: Support PP for Mistral Small 3.1 (#14254) 2025-12-02 13:04:14 +08:00
kun-llfl
106df4eac5 Fix mrope_positions size when req is retracted (#13700)
Signed-off-by: Kun(llfl) <i@imux.top>
Co-authored-by: Xuchun Shang <xuchun.shang@gmail.com>
2025-12-02 11:38:20 +08:00
Liangsheng Yin
427a19b643 Remove cargo config also in .zshenv (#14267) 2025-12-02 11:31:53 +08:00
Mick
1f930cd23d [diffusion] CI: add testcase-wise retry mechanism (#14261) 2025-12-02 11:06:12 +08:00
Kartik Ramesh
11ce05163d Fix NIXL exception message (#14172) 2025-12-02 10:39:45 +08:00
Simo Lin
cd4151abc7 [model-gateway] add audio and moderation in model card (#14263) 2025-12-01 18:36:38 -08:00
Stefan He
8fe8b63576 Revert "Try to remove wrong logic about max total token in spec decoding" (#14259) 2025-12-01 18:18:03 -08:00
Lianmin Zheng
8a7b1b8301 [Docs] Update CI docs (#14260) 2025-12-01 18:15:03 -08:00
Yuan Luo
26aebf83d3 [VLM] Support Piecewise CUDA Graph for Qwen3-Omni-MOE (#14222)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-12-02 10:12:10 +08:00
Mick
3ab8ae6847 [diffusion] fix: fix Flux.2 condition image resize (#14232)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-02 10:05:44 +08:00
Baizhou Zhang
03888b9de5 [Minor] Upgrade cutedsl version in Dockerfile (#13968) 2025-12-01 17:15:26 -08:00
Lianmin Zheng
796d82b107 [Auto Sync] Add max_total_num_tokens metric: Update scheduler_metrics_mixin.py, collector.py (20251202) (#14256)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Dan Zheng <dzheng@x.ai>
2025-12-01 16:34:34 -08:00
Lianmin Zheng
1da59e8304 [Auto Sync] optionally disable fake register in Update fp8_kernel.py (20251202) (#14255)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: gauravjain14 <41287729+gauravjain14@users.noreply.github.com>
2025-12-01 16:11:12 -08:00