Commit Graph
108 Commits
Author SHA1 Message Date
Lianmin ZhengandClaude Opus 4.6 46a392658e Refine RL & Post-Training description in README (#20877)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-18 12:43:42 -07:00
Yuwei AnandWenyao Gao 0abb9f4176 Piecewise Cuda Graph Docs (#19738)
Signed-off-by: yuweia <ayw.sirius19@gmail.com>
Co-authored-by: Wenyao Gao <wgao11@u.rochester.edu>
2026-03-03 11:51:17 +08:00
shuwenn 0c224c3c62 docs: Embed release lookup tool into Sphinx documentation site (#19264) 2026-02-24 11:11:27 -08:00
赵晨阳 e239f8aa85 Remove error dllm and diffusion doc in basic_useage (#19105) 2026-02-20 20:28:00 -08:00
qianyue76gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>JiaxinD
f06ab17a73 [diffusion] docs: consolidate diffusion documentation into docs (#18095)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: JiaxinD <djx2048@gmail.com>
2026-02-11 16:55:07 -08:00
AlexZhao赵海源gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>zhaochenyang20
3167bcc01c [Doc] Comprehensive Guide: Navigating DP, DPA, and SMG Best Practices (#18096)
Co-authored-by: 赵海源 <zhaohaiyuan@xiaohongshu.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>
2026-02-10 18:31:28 -08:00
Rishit ShivamRishitshivamRatish Pgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>Adarsh Shirawalmathzhaochenyang20
c850a8a41a [Docs] Add Falcon H1, Hunyuan-Large, Qwen3-Omni support and update Diffusion usage (#17888)
Co-authored-by: Rishitshivam <164783543+Rishitshivam@users.noreply.github.com>
Co-authored-by: Ratish P <114130421+Ratish1@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Adarsh Shirawalmath <114558126+adarshxs@users.noreply.github.com>
Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>
2026-02-06 13:17:51 -08:00
sglang-bot c971852ffc docs: move deepseek_ocr to popular model usage and add cookbook reference (#18120) 2026-02-02 05:45:41 -08:00
baonudesifeizhai 84ab611af8 model: support DeepSeek-OCR-2 (#17897) 2026-01-30 09:49:51 +08:00
Xiaoyu Zhang c08b54a575 [JIT kernel] Update jit_kernel cache and develop doc (#17842) 2026-01-28 15:09:47 +08:00
zijiexiaandJD dd97e1fe38 [Docs] Add RL documentation (#17663)
Co-authored-by: JD <jaedon.guo@gmail.com>
2026-01-26 12:16:54 -08:00
79ddc34c1c [Docs] Add new model evaluation docs (#17043)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
2026-01-19 16:35:03 -08:00
Adarsh Shirawalmath aab906a3d4 [docs] sync diffusion docs to main docs (#16932) 2026-01-12 14:49:55 +08:00
Yuan Luoandluoyuan.luo 9bd64d739b [VLM] Add doc for ViT CUDA Graph (#16343)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-01-04 10:05:23 +08:00
Lianmin Zheng c58a573a40 [Docs] Improve documentation index page (#16028) 2025-12-28 18:44:52 -08:00
Simo Lin 9665574937 [docs] major SGL Model Gateway documentation update (#15715) 2025-12-23 20:26:09 -08:00
Alison Shaogemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>Xinyuan Tong
31d48d7f6f Add Ollama-compatible API endpoints + Smart Router (#14376)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2025-12-16 20:43:38 -08:00
Tianyu Guo a9a2cdd8ec Add EPD disaggregation doc (#15224) 2025-12-16 11:22:05 +08:00
Shangming Caiandalpha-baby c8cf1cafdb [1/N] Update doc of Pipeline Parallelism (#14985)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: alpha-baby <fujianhao1997@qq.com>
2025-12-12 19:32:52 +08:00
76743a983e [DLLM] Add documentation for diffusion LLMs (#14358)
Co-authored-by: Tiwei Bie <tiwei.btw@antgroup.com>
Co-authored-by: Jinwei Yao <jinweiy@illinois.edu>
2025-12-11 20:29:51 -08:00
Tiance Wangwangtiancegemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
624725cb5e Move and update MindSpore docs, make it appear on the online documentation (#14861)
Co-authored-by: wangtiance <tiancew@qq.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-10 23:03:50 -08:00
88d1bab537 add doc for quantized kv cache (#14348)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
Co-authored-by: Ho-Ren (Jack) Chuang <horenchuang@bytedance.com>
2025-12-04 13:01:05 -08:00
dc1635023f [NPU][Doc] updated installation guide for Ascend NPU (#13585)
Co-authored-by: Howeee <15935120809@163.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2025-12-03 16:58:49 +03:00
b8zhongandBrayden Zhong 65c8568c4a sync attention, deepseek doc (#14335)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
2025-12-02 21:19:40 -08:00
b8zhongandBrayden Zhong e6420100ee sync attention doc and ep doc to doctree (#14257)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
2025-12-01 21:15:22 -08:00
Richard Chenandzhaochenyang20 4addb60274 Pull Request Instructions: RL and Training Framework Integrations (#14187)
Co-authored-by: zhaochenyang20 <zhaochenyang20@gmail.com>
2025-11-30 21:58:54 -08:00
Lianmin Zhengandsglang-bot 7e626d12b7 Update docs (#13391)
Co-authored-by: sglang-bot <sglangbot@gmail.com>
2025-11-16 19:36:33 -08:00
Adarsh ShirawalmathandUbuntu 583bb1804e [Docs] Add docs for Qwen3-VL image and video support (#12554)
Co-authored-by: Ubuntu <azureuser@athena.w2cgneqjjboeneyk2w5mje3jyf.bx.internal.cloudapp.net>
2025-11-10 12:16:04 +08:00
Teng Ma 1689c0e35f [Doc] fix miss index for production request trace (#12547) 2025-11-03 17:57:09 -08:00
Teng Ma 32438eba45 [Ckpt Engine] feat: new sglang entrypoint support for update (#12216) 2025-10-30 10:39:27 +08:00
Baizhou Zhang 8e987fa2a3 Update document index for DeepSeek-v32 docs (#12101) 2025-10-25 13:38:58 -07:00
thelongestusernameofallChengxing Xiegemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
49afb3d9d9 Fix(security): block unsafe pickle deserialization to mitigate CVE-2025-10164 (#11909)
Co-authored-by: Chengxing Xie <xiechengxing34@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-10-23 19:12:40 -07:00
Minglei Zhu 200a3c0bb1 [Documentation] add doc for deterministic inference (#11956) 2025-10-22 12:36:15 -05:00
b113c72e7a Init attention backend for Intel XPU (#10656)
Co-authored-by: guangyey <guangye.yu@intel.com>
Co-authored-by: DiweiSun <105627594+DiweiSun@users.noreply.github.com>
2025-10-21 11:41:28 +08:00
Lianmin Zheng 67e34c56d7 Fix install instructions and pyproject.tomls (#11781) 2025-10-18 01:08:01 -07:00
b8zhong 6bc503af73 [Doc] Update support matrix for attn and hybrid attn (#11293) 2025-10-14 22:43:11 -07:00
ykcombat f5754d1256 [Documentation][Configuration] Server args and documentation of PD-Multiplexing. (#11427) 2025-10-11 21:36:07 +08:00
ykwdandZhiqiang Xie 69efdd27bc [Doc] HiCache Design Documents (#11027)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2025-10-08 00:35:45 +08:00
Yi Zhang 760b788a58 add qwen3-next doc (#10327) 2025-09-11 14:29:11 -07:00
Huapeng Zhou 75ee00112d [Doc] Fix SGLang tool parser doc (#9886) 2025-09-04 21:52:53 +08:00
a85363c199 [docs] Instructions for bench_serving.py (#9071)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>
Co-authored-by: zhaochenyang20 <zhaochenyang20@gmail.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
2025-08-26 18:30:57 -07:00
Lianmin Zheng 2e8e7e353b Improve docs and developer guide (#9044) 2025-08-10 21:05:18 -07:00
Lianmin Zheng 2449a0afe2 Refactor the docs (#9031) 2025-08-10 19:49:45 -07:00
Lianmin Zheng 706bd69cc5 Clean up server_args.py to have a dedicated function for model specific adjustments (#8983) 2025-08-08 19:56:50 -07:00
44d600cd67 Support precomputed_embeddings for Llama 4 (#8156)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xiang (Kevin) Li <lik@nvidia.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2025-07-27 01:14:49 -07:00
Lianmin Zheng 0f218731e3 Do not run frontend_reasoning.ipynb to reduce the CI load (#7073) 2025-06-10 17:15:31 -07:00
Yudi Xue 14c18d25df Frontend language separate reasoning support (#6031) 2025-06-10 17:11:29 -07:00
Lianmin Zheng bb185b0e92 Update README.md (#7040) 2025-06-10 01:59:14 -07:00
Marc Sun 37f1547587 [FEAT] Add transformers backend support (#5929) 2025-06-03 21:05:29 -07:00
linzhuoandChayenne 7a0bbe6a64 update toc for doc and dockerfile code style format (#6450)
Co-authored-by: Chayenne <zhaochen20@outlook.com>
2025-05-27 13:05:11 +08:00