Commit Graph

6215 Commits

Author SHA1 Message Date
Jinwu
13bf565d60 [2/N]Support DeepSeek-R1 w4a8 low latency deepep (#8464)
Co-authored-by: Hank Han <hanhan7630@outlook.com>
Co-authored-by: Shangchuan Huang <2510421000@qq.com>
2025-10-24 17:41:16 -07:00
yinghui
e51046beaa perf: trtllm_mla attention backend spec decoding speedup w/ cuda graph (#12093) 2025-10-24 16:05:44 -07:00
Luca Lowndes
4eeeae1e75 Fix: Update blog link (#12071) 2025-10-24 15:26:23 -07:00
Minglei Zhu
f4b78d137c [1/2] deepseek deterministic: support deterministic inference for deepseek arch models on a single GPU (#12000) 2025-10-24 15:17:28 -07:00
Chang Su
4463e90dd2 [router][grpc] Remove gpt_oss parsers and remove _parser suffix in tool parser files (#12091) 2025-10-24 15:14:41 -07:00
Simo Lin
229f236da8 [router] migrate app context to builder pattern 2/n (#12089) 2025-10-24 15:04:32 -07:00
Simo Lin
5983e5bd1b [router] migrate app context to builder pattern 1/n (#12086) 2025-10-24 13:35:22 -07:00
Jonah Bernard
4b046a72d3 docs(server-arguments): add allowed options for each argument (#11560) 2025-10-24 11:49:20 -07:00
Simo Lin
770d63123c [router] fix ut router config init to use build pattern (#12084) 2025-10-24 11:16:31 -07:00
ishandhanani
14203432b4 fix(compile_utils, ep_moe): update environment variable and dtype check (#12034) 2025-10-24 11:00:12 -07:00
Keyang Ru
d7f0d88fa2 [router] implement response api get input item function and refactor input/output store (#11924) 2025-10-24 10:44:21 -07:00
Glen Liu
fc86b18b3e adjust dynamic vs static outputs comparison in test_lora_update.py (#11884) 2025-10-24 10:35:34 -07:00
0bfa394aff [Fix]: HiCache hasher failed when EAGLE mode enabled (#12025) 2025-10-24 23:53:13 +08:00
fzyzcjy
e04340bf48 Fix multi processing serializer bug (#11958) 2025-10-24 22:53:45 +08:00
Xiaoyu Zhang
8470133852 [b200] fix piecewise cuda graph launch bug (#12067) 2025-10-24 22:36:39 +08:00
Muqi Li
93ef9a094d [Profiler] expand '~' for torch_profiler_output_dir (#11999) 2025-10-24 17:20:46 +08:00
Muqi Li
b04cd3d487 Add 'gguf' to project dependencies (#12046) 2025-10-24 17:16:19 +08:00
Yuan Luo
7ef5d8afd4 Revise POINTSV15Chat model (#12049)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-10-24 17:09:45 +08:00
Qiaolin Yu
71d41212e4 Fix dpsk-r1-fp4 launching crash (#12063) 2025-10-24 17:04:50 +08:00
Xinyuan Tong
b9fb74f3bc fix: bench_serving ITL calculation when using spec-decoding (#12064) 2025-10-24 17:02:44 +08:00
ybyang
e15b63a182 [Fix] fix missing ipc_name of __getitem__ in some IO structs (#12053)
Signed-off-by: ybyang <ybyang7@iflytek.com>
2025-10-24 16:59:14 +08:00
Yuxuan Zhang
4060ed37cb Refactoring GLM-4.5 and GLM-4.5V related implementations (#11800) 2025-10-24 08:22:36 +00:00
fzyzcjy
2342605ef0 Tiny cleanup send_single (#12056) 2025-10-23 23:53:42 -07:00
Simo Lin
dbf17a8313 [router] Add mTLS Support for Router-to-Worker Communication (#12019)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
2025-10-23 23:14:09 -07:00
fzyzcjy
0f0c430e93 Install numactl in Dockerfile for GH200/GB200/GB300 (#11853) 2025-10-23 21:39:10 -07:00
Rain Jiang
8e797a47f0 fix: the hardcode hf repo name comparison for deepseek-ocr (#12031)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-10-23 21:37:56 -07:00
Zaili Wang
aa3003f116 Add gguf dependency for cpu/xpu (#12041) 2025-10-23 21:13:17 -07:00
Yongfei Xu
4793ec7d1a Opt MHA chunked prefix: merge prefix and extend kv cache to run mha once (#10953) 2025-10-23 20:58:10 -07:00
Zaili Wang
92009bd28e fix: fix MMMU loading issue (#11759)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-10-23 20:21:38 -07:00
Baizhou Zhang
4ef981e2b6 Revert "[Fix] Fix lint to pass CI" (#12042) 2025-10-23 19:44:58 -07:00
Baizhou Zhang
69ed8b67a8 [Fix] Fix lint to pass CI (#12037) 2025-10-23 19:39:38 -07:00
narutolhy
1801cd199f support more model in piecewise cuda graph (#11745) 2025-10-24 10:31:39 +08:00
Lianmin Zheng
ffc722a690 Revert "lang: support direct video inference" (#12038) 2025-10-23 19:21:31 -07:00
thelongestusernameofall
49afb3d9d9 Fix(security): block unsafe pickle deserialization to mitigate CVE-2025-10164 (#11909)
Co-authored-by: Chengxing Xie <xiechengxing34@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-10-23 19:12:40 -07:00
b8zhong
f80371ff8c Use flashinfer_trtllm moe runner backend to gain around 10% perf on b200 fp8 dpsk (#11816) 2025-10-23 19:12:15 -07:00
Jonah Bernard
62eff37ba1 Refactor Triton-kernel MoE runner integration (#11795) 2025-10-23 18:47:28 -07:00
b8zhong
47e12e082e Enable Llama 4 + TRTLLM MHA (#12003) 2025-10-23 18:22:58 -07:00
Mick
823b442945 lang: support direct video inference (#9936)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2025-10-23 18:12:39 -07:00
Fan Yin
14a4d80e57 [8/n] decouple quantization impl from vllm dependency - gguf srt (#11964)
Co-authored-by: Peng Zhang <zhuangsen.zp@antgroup.com>
2025-10-23 18:12:00 -07:00
sglang-bot
1053e1be17 chore: bump SGLang version to 0.5.4 (#12027)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2025-10-23 18:01:40 -07:00
nvjullin
9a71500cfb Fixed aarch64 flash-mla (#12009) 2025-10-23 17:47:04 -07:00
Simo Lin
6d6e24bcc4 [router] Add builder pattern for RouterConfig with zero duplication (#12030) 2025-10-23 16:46:10 -07:00
Kangyan-Zhou
2c057fbfa8 Update Github action title for kernel build (#12029) 2025-10-23 13:39:40 -07:00
Roger Young
dbd9435dc1 Fix mamba radix cache eviction logic in alloc_req_slots (#11616)
Signed-off-by: rogeryoungh <rogeryoungh@foxmail.com>
2025-10-23 13:07:43 -07:00
b8zhong
8ae9d4bb41 Revert "[ROCm] Remove vLLM rope dependency & use AITER impl" (#12028) 2025-10-23 12:42:59 -07:00
Nicolas Castet
1c304aa9bc Log iteration # for prefill and decode (#9366) 2025-10-23 12:28:03 -07:00
Mick
770529a731 model: support deepseek-ocr (#11891)
Co-authored-by: yhyang201 <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Shi Shuai <126407087+shuaills@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2025-10-24 03:15:17 +08:00
ErvinXie
39c237f02c Add AWQ quantization support for NPU. (#10158)
Co-authored-by: Alisehen <814073252@qq.com>
Co-authored-by: Yaochen Han <48639761+Alisehen@users.noreply.github.com>
Co-authored-by: Zhengda Qin <zhengdqin@gmail.com>
2025-10-23 12:08:05 -07:00
Chang Su
28b8a4064d [router][CI] Clean up imports and prints statements in sgl-router/py_test (#12024) 2025-10-23 11:56:57 -07:00
Mick
8bd26dd4e6 ci: fix night-ci with push retry mechanism (#11765) 2025-10-23 11:31:05 -07:00