Yuhong Guo
|
4d84f886e7
|
Refactor --debug-tensor-dump-layers to list (#12691)
|
2025-11-05 03:30:01 -08:00 |
|
Trevor Morris
|
dbcf85b7f0
|
Add --speculative-moe-runner-backend server arg (#10183)
|
2025-11-04 00:20:56 -08:00 |
|
Lianmin Zheng
|
f600866a44
|
Improve the metrics for PD (#12580)
Co-authored-by: Kan Wu <wukanustc@gmail.com>
Co-authored-by: cctry <shiyang@x.ai>
|
2025-11-03 22:10:57 -08:00 |
|
Lianmin Zheng
|
7a21d8b276
|
Reduce the overhead of nccl symmetric memory (#12524)
Co-authored-by: Nicolas Castet <ncastet@nvidia.com>
|
2025-11-03 11:56:27 -08:00 |
|
Hanming Lu
|
66fb9b1307
|
[ServerArgs] allow --mamba-ssm-dtype extend (#12481)
|
2025-11-02 11:50:04 -08:00 |
|
kousakawang
|
7efd8b3d1f
|
[FEAT] Shared mem pool based cuda ipc for multi-modal data transport (#11917)
Co-authored-by: kousakawang <wanghanpei@bytedance.com>
Co-authored-by: Yuan Luo <4908075+yuan-luo@users.noreply.github.com>
|
2025-11-02 16:46:37 +08:00 |
|
Ho-Ren (Jack) Chuang
|
76196b3cbf
|
feat: Add FP4 (E2M1) KV Cache Support with Quantization Utilities for MLA (#10078)
Signed-off-by: Ho-Ren (Jack) Chuang <horenchuang@bytedance.com>
Co-authored-by: Yichen Wang <yichen.wang@bytedance.com>
|
2025-11-01 22:24:58 -07:00 |
|
Xinyuan Tong
|
d2a8f71c2f
|
[feat] Add SGLANG_TOOL_STRICT_LEVEL for tool-call behavior control (#12423)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-11-01 13:15:02 -07:00 |
|
Ke Bao
|
a4bf5c6ad2
|
Support Kimi Linear (#12469)
Co-authored-by: yizhang2077 <1109276519@qq.com>
|
2025-10-31 14:03:35 -07:00 |
|
Yuan Luo
|
c30ebb9300
|
[VLM] Optimize async mm data process mechanism (#12066)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-11-01 01:24:53 +08:00 |
|
b8zhong
|
a076ec1a7a
|
Revert "fix llama4 kv cache layout" (#12437)
|
2025-10-30 22:33:37 -07:00 |
|
Yuhong Guo
|
ab95d35fcb
|
feat: Add Non-intrusive Tensor Dumping for Model Inference (#10566)
|
2025-10-31 12:04:48 +08:00 |
|
Lianmin Zheng
|
6a63a9852e
|
minor code sync (#12403)
|
2025-10-30 12:49:31 -07:00 |
|
b8zhong
|
bacb3825fe
|
fix: llama 4 + trtllm gen + fp8 kv cache incompatibility (#12347)
|
2025-10-29 11:31:02 -07:00 |
|
Rain H
|
750940ae36
|
Eagle3 DP attention for Qwen3 MoE (#12002)
|
2025-10-29 20:25:17 +08:00 |
|
Baizhou Zhang
|
685c06451f
|
[ci] Try fixing broken CIs (#12317)
|
2025-10-29 01:13:51 -07:00 |
|
b8zhong
|
c143f416ce
|
fix: Llama 4 BF16 load on Blackwell (#12308)
|
2025-10-28 18:59:01 -07:00 |
|
b8zhong
|
77225d602a
|
Use Flashinfer TRT-LLM as Llama 4 compatible MoE backend (#11928)
|
2025-10-28 10:39:43 -07:00 |
|
Trevor Morris
|
fdd00295b5
|
Fix 'BypassedTopKOutput' object has no attribute 'topk_weights' for DeepEP (#12231)
|
2025-10-28 09:28:25 -07:00 |
|
Feng Su
|
ea96106000
|
[Feature] Sglang Tracing: Fine-Grained Tracking for Request Latency - Part 2 (#10804)
Signed-off-by: Feng Su <sufeng@linux.alibaba.com>
|
2025-10-28 01:25:46 -07:00 |
|
hlu1
|
81a632ace6
|
[DeepseekV32] Enable flashmla_prefill kernel with fp8 kvcache (#11655)
Signed-off-by: Hao Lu <14827759+hlu1@users.noreply.github.com>
|
2025-10-27 23:11:48 -07:00 |
|
Lifu Huang
|
ce832d7034
|
Add env var to control custom Triton kernel cache and set CSGMV as default backend. (#12176)
|
2025-10-27 17:49:32 -07:00 |
|
weiliang
|
88596739a4
|
Support running FP4 Deepseek on SM120. (#11708)
|
2025-10-27 17:37:49 -07:00 |
|
Weiwei
|
caa4819bfc
|
Add support for AutoRound quantized models (#10153)
|
2025-10-27 18:17:29 +08:00 |
|
Liangsheng Yin
|
8491c794ad
|
[misc] depdencies & enviroment flag (#12113)
|
2025-10-26 14:52:35 +08:00 |
|
Lianmin Zheng
|
7b36c47b3b
|
Clean up attention backend selection code & Other minor rename (#12136)
|
2025-10-25 23:50:12 -07:00 |
|
Lianmin Zheng
|
8e70064c37
|
Clean up server launch code and multi tokenizer (#12132)
|
2025-10-25 16:40:27 -07:00 |
|
Lianmin Zheng
|
4caca1ba04
|
Clean up server args & Add CI scripts (#12124)
|
2025-10-25 11:53:57 -07:00 |
|
fzyzcjy
|
20bd2271e2
|
Support true on-policy (#12058)
|
2025-10-25 10:23:42 +08:00 |
|
Minglei Zhu
|
f4b78d137c
|
[1/2] deepseek deterministic: support deterministic inference for deepseek arch models on a single GPU (#12000)
|
2025-10-24 15:17:28 -07:00 |
|
Jonah Bernard
|
4b046a72d3
|
docs(server-arguments): add allowed options for each argument (#11560)
|
2025-10-24 11:49:20 -07:00 |
|
b8zhong
|
f80371ff8c
|
Use flashinfer_trtllm moe runner backend to gain around 10% perf on b200 fp8 dpsk (#11816)
|
2025-10-23 19:12:15 -07:00 |
|
b8zhong
|
47e12e082e
|
Enable Llama 4 + TRTLLM MHA (#12003)
|
2025-10-23 18:22:58 -07:00 |
|
Teng Ma
|
96a5e4dd79
|
[Feature] Support loading weights from ckpt engine worker (#11755)
Signed-off-by: Yang Kaiyong <yangkaiyong.yky@antgroup.com>
Signed-off-by: Cruz Zhao <CruzZhao@linux.alibaba.com>
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: Yang Kaiyong <yangkaiyong.yky@antgroup.com>
Co-authored-by: Cruz Zhao <CruzZhao@linux.alibaba.com>
Co-authored-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2025-10-23 09:23:30 -07:00 |
|
Jue WANG
|
138ff23187
|
Allow to disable batch decoding. (#11944)
|
2025-10-22 21:57:12 -07:00 |
|
Hank Han
|
904655c5fd
|
[2/N] Added the core structure of elastic EP and the eplb algorithm with faulty rank (#10606)
Co-authored-by: Xun Sun <UNIDY2002@outlook.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2025-10-22 01:13:31 -07:00 |
|
Zhiyu
|
80b2b3207a
|
Enable native ModelOpt quantization support (3/3) (#10154)
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
|
2025-10-21 21:44:29 -07:00 |
|
Baizhou Zhang
|
ef4a8097b8
|
Rename flashmla kernel options of nsa backend for better readability (#11876)
|
2025-10-21 13:14:16 -07:00 |
|
Atream
|
7e6191c098
|
init support for KTransformers Heterogeneous Computing (#11487)
Co-authored-by: Jianwei Dong <1913953267@qq.com>
|
2025-10-21 00:17:02 -07:00 |
|
Meng, Hengyu
|
b113c72e7a
|
Init attention backend for Intel XPU (#10656)
Co-authored-by: guangyey <guangye.yu@intel.com>
Co-authored-by: DiweiSun <105627594+DiweiSun@users.noreply.github.com>
|
2025-10-21 11:41:28 +08:00 |
|
Lianmin Zheng
|
43ad05907c
|
[Auto Sync] Update scheduler.py, server_args.py (20251020) (#11875)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Kan Wu <wukanustc@gmail.com>
|
2025-10-20 17:41:19 -07:00 |
|
ishandhanani
|
296f689242
|
fix(server_args): handle tokenizer init conflicts (#11776)
|
2025-10-20 00:27:19 -07:00 |
|
Shane A
|
d383e6616e
|
[Model] Add Olmo 3 model support (#11396)
|
2025-10-19 23:59:16 -07:00 |
|
Johnny
|
252dc4e112
|
[NVIDIA] FA3/FA4 Fix (#11606)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-10-19 17:10:10 -07:00 |
|
Baizhou Zhang
|
cbb5fc2edc
|
[CI] Add CI test for DeepSeek V3.2 MTP (#11835)
|
2025-10-19 17:00:25 -07:00 |
|
Stefan He
|
4fff1ec1d9
|
Deterministic Mode: Add 1-stage triton kernel for prefill (#11147)
Co-authored-by: Minglei Zhu <mingleizhu1122@gmail.com>
Co-authored-by: Binyao Jiang <bijiang@linkedin.com>
|
2025-10-20 01:47:36 +08:00 |
|
b8zhong
|
f9a7d9b3dc
|
support server arg override KV cache to bf16 to avoid slow cases (#11749)
|
2025-10-19 02:49:48 +08:00 |
|
Lianmin Zheng
|
67e34c56d7
|
Fix install instructions and pyproject.tomls (#11781)
|
2025-10-18 01:08:01 -07:00 |
|
Yuwei An
|
1d726528f7
|
Eager Compiler for Torch Compile (#11803)
Signed-off-by: Oasis-Git <ayw.sirius19@gmail.com>
|
2025-10-18 15:18:52 +08:00 |
|
Minglei Zhu
|
f4488e9dd9
|
set default attention backend for deterministic inference (#11801)
|
2025-10-18 00:01:24 -07:00 |
|