Commit Graph
1261 Commits
Author SHA1 Message Date
shuwenn bc2405e6c1 feat: support release lookup (#18450) 2026-02-13 10:47:02 +08:00
danielafrimi e422bcaed8 [Mamba] Add float16 support for SSM cache dtype (#18444) 2026-02-12 11:27:47 +08:00
fy 123f57b84b update glm5 readme on npu (#18657) 2026-02-12 10:37:12 +08:00
liupeng374 c34832c02c glm5 md (#18655) 2026-02-12 10:11:59 +08:00
qianyue76gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>JiaxinD
f06ab17a73 [diffusion] docs: consolidate diffusion documentation into docs (#18095)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: JiaxinD <djx2048@gmail.com>
2026-02-11 16:55:07 -08:00
Baizhou Zhang 947927bdb5 [V3.2] Change default CP token split method to --round-robin-split (#18613) 2026-02-11 20:14:35 +08:00
赵晨阳 a2c38f7796 Enhance SMG guide with RL rollout systems benefits (#18588) 2026-02-10 20:20:45 -08:00
AlexZhao赵海源gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>zhaochenyang20
3167bcc01c [Doc] Comprehensive Guide: Navigating DP, DPA, and SMG Best Practices (#18096)
Co-authored-by: 赵海源 <zhaohaiyuan@xiaohongshu.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>
2026-02-10 18:31:28 -08:00
husf 573ff55814 [NPU][docs]fix bug about hyperlink for best practice for ascend npu (#18561) 2026-02-10 20:03:28 +03:00
Hexq0210 d0d387dea1 [NPU] update npu doc (#18474) 2026-02-10 16:59:13 +03:00
husf 99101ce30b [NPU][docs] improve docs for Best Practice on Ascend NPU (#18360) 2026-02-10 16:52:27 +03:00
Zack YuBrayden Zhonggemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
54589a2f2d docs: expand and update modelopt documentation (#18479)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-09 23:09:52 +00:00
brimon ddbcfbaaab feature: support bidirectional attention for Gemma-3 (#10707) 2026-02-09 23:17:45 +08:00
Junlin ZhouandTiwei Bie 14652243bd [DLLM] Add JointThreshold algorithm for joint M2T and T2T decoding (#18171)
Signed-off-by: Junlin Zhou <zhoujunlin.zjl@antgroup.com>
Co-authored-by: Tiwei Bie <tiwei.btw@antgroup.com>
2026-02-09 14:20:45 +08:00
Mohammad Miadh Angkad fddef76619 [Doc] Fix outdated --fp4-gemm-backend documentation (#18350) 2026-02-07 20:42:47 +08:00
Mohammad Miadh AngkadandBaizhou Zhang c47c2f9466 [Doc] Update CUDA 13 install guide to install torch first (#18404)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-02-07 18:04:37 +08:00
Hexq0210 e834b85ab6 [NPU] update npu doc (#18344) 2026-02-07 16:38:05 +08:00
Rishit ShivamRishitshivamRatish Pgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>Adarsh Shirawalmathzhaochenyang20
c850a8a41a [Docs] Add Falcon H1, Hunyuan-Large, Qwen3-Omni support and update Diffusion usage (#17888)
Co-authored-by: Rishitshivam <164783543+Rishitshivam@users.noreply.github.com>
Co-authored-by: Ratish P <114130421+Ratish1@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Adarsh Shirawalmath <114558126+adarshxs@users.noreply.github.com>
Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>
2026-02-06 13:17:51 -08:00
amote-i 92b8bd6833 fix npu best practice (#18330) 2026-02-05 21:14:46 -05:00
shuwenn ef1d0ea885 [Doc] add a summary section for spec decode document (#18323) 2026-02-05 16:34:31 -05:00
shuwenn 8b21dd4b77 [Doc] refine spec decode docs for SpecV2/STANDALONE/NGRAM (#18321) 2026-02-05 15:12:33 -05:00
Kun Lin e616d35847 Support Markdown/Notebook-Friendly Documentation Export for Downstream Integration(convert rat files to md files and save) (#18278) 2026-02-04 19:59:40 -08:00
rinbaro de6a03260f [docs] fix misspellings & typos (#18276) 2026-02-05 03:35:29 +00:00
Teng MaandShangming Cai c8212b9fac [PD] doc: Document SGLANG_MOONCAKE_CUSTOM_MEM_POOL and supported values (#18259)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-02-05 03:03:33 +00:00
Baizhou Zhang d279520ba5 [DeepGemm] Add a flag for fast warmup (#18111) 2026-02-04 14:12:13 +08:00
Kun Lin 669a9bd180 Support Markdown/Notebook-Friendly Documentation Export for Downstream Integration (copy all markdown and rst files) (#18223) 2026-02-03 20:53:17 -08:00
Viacheslav 74f716dbd7 Gigachat 3 tool parser and tests (#14765) 2026-02-02 22:28:34 -08:00
Kun Linand赵晨阳 f032c4f3d6 Support Markdown/Notebook-Friendly Documentation Export for Downstream Integration (#18131)
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
2026-02-02 21:43:20 -08:00
sglang-bot c971852ffc docs: move deepseek_ocr to popular model usage and add cookbook reference (#18120) 2026-02-02 05:45:41 -08:00
Yuan Luoandluoyuan.luo afebb7ab78 Optimize custom-all-reduce (#17674)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-02-01 18:59:31 +08:00
Baizhou Zhang c7d53fa26a Set torch url index in pyproject.toml (#16802) 2026-02-01 13:23:52 +08:00
Zaili Wang 97593c9f41 [CPU] toml file update (#17861) 2026-01-31 13:16:06 -08:00
husf c52578c7fd 【docs】【NPU】Update Expert Parallelism docs for Ascend NPU (#17940) 2026-01-30 23:42:11 -05:00
Tiance Wangandwangtiance f6a4ff718f doc update for CANN version (#18014)
Co-authored-by: wangtiance <tiancew@qq.com>
2026-01-30 21:10:23 -05:00
amote-i 9c2a468e2c update ascend docs (#17987) 2026-01-30 18:03:55 +08:00
Wenchen Loandb8zhong 046b29be16 GPTJForCausalLM Support (#7839)
Co-authored-by: b8zhong <b8zhong@uwaterloo.ca>
2026-01-29 21:00:04 -08:00
baonudesifeizhai 84ab611af8 model: support DeepSeek-OCR-2 (#17897) 2026-01-30 09:49:51 +08:00
StonyPortandqiuxuan.lzw 2b3408ff14 feat: add forward timeout (#17831)
Co-authored-by: qiuxuan.lzw <qiuxuan.lzw@alibaba-inc.com>
2026-01-30 08:52:29 +08:00
Ziang Li 3c9cc44ff5 Add mxfp8 support for online quantization, Triton dense linear, and CUTLASS MoE (#17449) 2026-01-29 21:33:57 +08:00
RoyWangandroywang 30adf78f82 [diffusion]: align sglang diffusion AMD pyproject_other.toml diffusion dependency with pyproject.toml (#16225)
Co-authored-by: roywang <roywang@amd.com>
2026-01-29 01:50:57 -08:00
amote-i 1b22f2ee1c update ascend docs (#17741) 2026-01-29 11:43:48 +08:00
Joe Redmond 0ff0d181ca feat: add custom request header logging (#17786) 2026-01-28 19:33:08 -08:00
f1384f5293 Integration mori backend for EP a2a data communication (#17012)
Co-authored-by: Duyi-Wang <duyi.wang@amd.com>
Co-authored-by: billishyahao <bill.he@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
2026-01-28 19:07:34 -08:00
Артем Савкинandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> b77b0ffd60 [NPU] NZ for non-quantized MOE, Qwen3 MOE double memory consumption fix (#15904)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-29 00:55:08 +08:00
Xiaoyu Zhang c08b54a575 [JIT kernel] Update jit_kernel cache and develop doc (#17842) 2026-01-28 15:09:47 +08:00
Kangyan-Zhou c0b4dd68a2 Add a performance dashboard server and frontend for nightly CUDA tests (#17725) 2026-01-27 22:22:33 -08:00
Hubert Lu 93423ff780 [AMD] Deprecate ROCm 6.3 artifacts and standardize gfx942 on ROCm 7 (#17785) 2026-01-27 15:58:49 -08:00
Baizhou Zhang 1d942e4eef [DeepSeek] Update tests and document for DeepSeek V3.2 NVFP4 checkpoint (#17657) 2026-01-27 22:10:57 +08:00
Baizhou Zhang 832c756549 [Doc] Tiny update description on torch compile (#17819) 2026-01-27 18:59:04 +08:00
7106f6c8e1 [GLM-OCR] Support GLM-OCR Model (#17582)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-01-26 22:24:00 -08:00