Commit Graph

1015 Commits

Author SHA1 Message Date
Hubert Lu
93423ff780 [AMD] Deprecate ROCm 6.3 artifacts and standardize gfx942 on ROCm 7 (#17785) 2026-01-27 15:58:49 -08:00
Baizhou Zhang
1d942e4eef [DeepSeek] Update tests and document for DeepSeek V3.2 NVFP4 checkpoint (#17657) 2026-01-27 22:10:57 +08:00
Baizhou Zhang
832c756549 [Doc] Tiny update description on torch compile (#17819) 2026-01-27 18:59:04 +08:00
Yuxuan Zhang
7106f6c8e1 [GLM-OCR] Support GLM-OCR Model (#17582)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-01-26 22:24:00 -08:00
Hubert Lu
df42f4d386 [AMD] Update dsv3.2 AMD GPU docs and unify ROCm TileLang build (#17783)
Co-authored-by: wufann <715544327@qq.com>
2026-01-26 21:10:32 -08:00
shuwenn
fd3b179ffd [HiCache][HA 1/N] Support HiCache storage runtime attach/detach (#15892) 2026-01-26 19:33:19 -08:00
zijiexia
dd97e1fe38 [Docs] Add RL documentation (#17663)
Co-authored-by: JD <jaedon.guo@gmail.com>
2026-01-26 12:16:54 -08:00
Makcum888e
bba6e38ff8 [NPU] Split pyproject npu from pyproject other (#17641) 2026-01-26 09:45:44 -08:00
CSWYF3634076
1a19b3987d [Model] Add Ernie4.5 VL model support (#15679)
Signed-off-by: CSWYF3634076 <wangyafeng@baidu.com>
Signed-off-by: wangyafeng <wangyafeng@baidu.com>
2026-01-25 22:36:29 -08:00
zackyoray
d275d47973 [NIXL] Add custom NIXL backend selection for KVManager (#17146)
Signed-off-by: Yoray Zack <yorayz@nvidia.com>
2026-01-26 14:35:38 +08:00
Trevor Morris
2c2c4e446b [NVIDIA] Add flashinfer all-to-all MOE dispatcher (#14668) 2026-01-24 22:59:55 +08:00
Glen Liu
a6280b2a23 add documentation example for LoRA overlap loading and cleanup unused function (#17464) 2026-01-24 15:33:16 +08:00
R0CKSTAR
628ab5d57b [MUSA][2/N] sgl-kernel build (#17053)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
2026-01-23 14:41:47 -08:00
Mansoor
bdaa3de075 Add return routed experts to the completions and chat/completions endpoints (#17434) 2026-01-23 12:12:36 -08:00
Yi Zhong
08fcda2f63 add the fa4 mm backend and varlen func (#13539)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-01-23 23:12:06 +08:00
Nicolas Castet
48e9daadff Support symmetric memory pre-allocation to avoid fragmentation (#17089) 2026-01-23 17:57:04 +08:00
Hexq0210
69470dbc1f [NPU] update doc for Ascend NPU (#17621) 2026-01-23 15:21:00 +08:00
Lianmin Zheng
56e6652d1d Lazy import torchao (#17626) 2026-01-22 22:04:51 -08:00
amote-i
0fec8820d1 update dependence docs of npu (#17573) 2026-01-23 09:13:36 +08:00
hxie
13f88045b3 configuration file support and nixl integration augmentation for hicache-storage-backend-extra-config (#16602) 2026-01-22 14:31:48 -08:00
Zaili Wang
672eb37534 [CPU][Fix CI] Solidate torch version for sgl-kernel-cpu and fix device orientation error (#17460) 2026-01-22 14:04:50 +08:00
Lianmin Zheng
b74a57a8d9 [Auto Sync] Update detokenizer_manager.py, io_struct.py, mu... (20260120) (#17442)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Wangfan Fu <wangfan@x.ai>
2026-01-21 14:48:32 -08:00
Lingjun Wen
cf89351691 [new-model] Add support for Cohere2ForCausalLM behind Command-A and Command-R Models (#16927) 2026-01-21 12:28:33 -08:00
Yi Zhong
458fe5a337 [docs] Show user the fastAPI docs available (#17510)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
2026-01-21 14:26:25 +00:00
amote-i
0a9099e137 update ascend docs (#17457) 2026-01-21 15:36:26 +08:00
khalilzhk
aca354bcb3 [NPU] remove features supported on Ascend NPU (#17455) 2026-01-21 11:00:04 +08:00
zijiexia
4ecd9afde9 [Docs] Rename SGLang Router to SGLang Model Gateway (#17436) 2026-01-20 12:31:10 -08:00
Ruihuan He
eb38d64413 [Docs] Explain CUDA attention backend choices and aiter FP8 KV cache support (#17428) 2026-01-20 11:10:36 -08:00
Baizhou Zhang
6ea491e439 Overlap shared experts with deepep dispatch for single batch overlap on Blackwell (#17289) 2026-01-21 02:56:55 +08:00
Baizhou Zhang
55c616427d Add flag that enables NCCL mlp sync batch for overlap scheduler (#17288) 2026-01-20 23:06:55 +08:00
Shangming Cai
23d765d1f2 [Doc] Update pipeline parallelism documentation (#17414) 2026-01-20 20:10:56 +08:00
amote-i
603f386c6b update docs of Ascend plateform (#17358) 2026-01-20 19:31:23 +08:00
Hexq0210
1b192cf198 [NPU] Update NPU doc for model and features supported (#17385) 2026-01-20 12:28:06 +08:00
zijiexia
9f8b79f16f [Docs] Fix formatting in Evaluating New Models with SGLang (#17376) 2026-01-19 18:22:30 -08:00
zijiexia
79ddc34c1c [Docs] Add new model evaluation docs (#17043)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
2026-01-19 16:35:03 -08:00
b8zhong
f374623fa9 [Refactor] Set fp4-gemm-backend=auto on SM100 and rename fp4-gemm-backend with flashinfer_ prefix (#17309) 2026-01-19 20:09:07 +08:00
Glen Liu
ad1b4e4728 [Feature] overlap LoRA weight loading with compute (#15512) 2026-01-19 10:43:17 +08:00
Xinyuan Tong
2069050d3f fix: Handle multiple named chat templates in HuggingFace tokenizers (#17236)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-01-18 17:20:04 +08:00
Mohammad Miadh Angkad
088758c1c1 [Tiny] Improve docs (#17264) 2026-01-18 14:57:01 +08:00
zijiexia
166396ca4c [Docs] minor update on ep docs (#17242) 2026-01-16 18:20:47 -08:00
Yi Zhong
ec9b48ea96 Add olmo3 in supported docs (#13672)
Signed-off-by: Vincent Zhong <207368749+vincentzed@users.noreply.github.com>
2026-01-16 12:18:16 -05:00
Mohammad Miadh Angkad
c771933dc5 [Doc] Tiny docs update for CUDA 13 (#17200) 2026-01-16 20:53:36 +08:00
Raghav Ravishankar
daea51385d Add AFMoE model implementation (#13216) 2026-01-16 20:35:42 +08:00
StonyPort
3355b6e21b feat: add request queued timeout (#17143)
Co-authored-by: qiuxuan.lzw <qiuxuan.lzw@alibaba-inc.com>
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
2026-01-16 17:55:09 +08:00
R0CKSTAR
a1dd3d48ac [diffusion] hardware: support diffusion (single GPU, 3/N) (#17105)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
2026-01-16 17:01:09 +08:00
Adarsh Shirawalmath
7c39ea68f3 [diffusion] model: support flux Klein (#17173) 2026-01-16 16:16:17 +08:00
shuwenn
8ec160ed46 feature: support uvicorn access log filter(disable logging /metrics) (#15513) 2026-01-15 20:00:06 -08:00
Baizhou Zhang
8b99af9af8 [Doc] Tiny update Cuda 13 environment instructions (#17174) 2026-01-16 06:12:26 +08:00
b8zhong
3d72944fb8 [Doc] Add tip on how to use Spec V2 (#15455) 2026-01-16 05:30:18 +08:00
Yi Zhong
7dde3438e2 Show how to use cu13 image with B300 (#17170)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
2026-01-15 15:25:20 -05:00