ThomasX
|
95bacb8862
|
fix: preserve HiCache load-back allocator headroom
|
2026-05-08 00:18:33 +08:00 |
|
 
|
bb737d7a82
|
Support Qwen3 MoE context parallel (#18233)
Co-authored-by: Shunkang <182541032+Shunkangz@users.noreply.github.co>
Co-authored-by: Jiying Dong <87510204+dongjiyingdjy@users.noreply.github.com>
|
2026-03-22 01:27:20 -07:00 |
|
hzh0425
|
c43d495dd5
|
[RadixTree][9/N Refactor]: Support unified init_load_back params (#20590)
|
2026-03-18 11:19:52 +08:00 |
|
pansicheng
|
97b2a89334
|
[RadixTree][8/N Refactor]: unify lock interface (#20330)
|
2026-03-16 11:49:51 +08:00 |
|
Michelle Wu
|
b7f13a7b73
|
[NPU] bugs fix for Deepseek models (#19544)
|
2026-02-28 17:26:15 +08:00 |
|
 chenxu214andsglang-npu-bot
|
5f07ff9271
|
Added the prefill delayer policy: The prefill deplay range is expanded. (#17456)
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
|
2026-02-28 08:56:49 +08:00 |
|
Yilong Zhao
|
07da4bed7b
|
[cache] add conservative estimation (#19482)
|
2026-02-27 18:14:46 +08:00 |
|
JD
|
191d354f53
|
fix double-free kv cache for requests that have already finished and been freed during preemption (#18694)
|
2026-02-13 13:17:44 -08:00 |
|
Zehuan Li
|
26f2b3798d
|
[DLLM] Basic dLLM scheduling strategy and implementation (#17484)
Signed-off-by: Zehuan Li <lizehuan.lzh@antgroup.com>
|
2026-02-10 16:54:15 +08:00 |
|
Jacob Gordon
|
15b511771d
|
refactor(codespell): corrects typos covered up by whitelist (#17601)
|
2026-01-22 11:45:47 -08:00 |
|
zhangheng
|
f33022d039
|
[RadixTree][3/N Refactor]:Support unified insert/evict params (#17401)
|
2026-01-22 17:36:31 +08:00 |
|
 
|
20b0523eca
|
[RadixTree][1/N Refactor]: Support unified match_prefix params (#17142)
Co-authored-by: yizhang2077 <1109276519@qq.com>
Co-authored-by: pansicheng <sicheng.pan.chn@gmail.com>
|
2026-01-19 22:39:40 +08:00 |
|
Zehuan Li
|
d2c863878c
|
[DLLM] Implement initial dynamic batching for diffusion LLM (#14883)
|
2026-01-17 16:48:15 +08:00 |
|
Ke Bao
|
7f8a58fffb
|
Refactor prefix cache type checking (#17028)
|
2026-01-15 11:28:13 +08:00 |
|
fzyzcjy
|
1f0ea4f958
|
Add routing key based schedule policy (#16840)
|
2026-01-10 10:16:23 +08:00 |
|
Ke Bao
|
76bc07a335
|
Move swa memory pool to a seperate file (#16347)
|
2026-01-04 22:39:30 +08:00 |
|
fzyzcjy
|
e797f0c570
|
Support offline generation scenario for prefill delayer (#16363)
|
2026-01-04 11:05:10 +08:00 |
|
Yongfei Xu
|
0d244116d2
|
[DeepSeek v3.2] opt Context Parallelism: support fused moe, multi batch and fp8 kvcache (#13959)
|
2026-01-02 23:49:14 +08:00 |
|
Ke Bao
|
c483a5f45f
|
Tiny adjust hybrid swa handling (#16292)
|
2026-01-02 23:11:23 +08:00 |
|
Cheng Wan
|
60f1ca6925
|
Refactor: Moving extend_logprob_start_len calculation out of prepare_for_extend (#16105)
|
2025-12-30 12:38:33 +08:00 |
|
  
|
7759716786
|
bugfix[schedule]: Refactor sort method and add related UT (#13576)
Co-authored-by: Yuxuan Wei <w1300012920@pku.edu.cn>
Co-authored-by: Yuxuan Wei <w1300012920@gmail.com>
Co-authored-by: Yuxuan Wei🚚 <yuxwei@microsoft.com>
|
2025-12-23 03:13:57 +08:00 |
|
 
|
45a959d3e9
|
[PP] Add pp support for Qwen3-VL (#12333)
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Signed-off-by: Kun(llfl) <i@imux.top>
Signed-off-by: Kun(llfl) <llfl@linux.alibaba.com>
Co-authored-by: kun-llfl <i@imux.top>
Co-authored-by: Kun(llfl) <llfl@linux.alibaba.com>
|
2025-12-17 16:03:58 +08:00 |
|
Hanming Lu
|
e61dabf5e4
|
[Qwen3-next] support mamba radix cache for overlap scheduler (#14792)
|
2025-12-14 18:54:16 -08:00 |
|
fzyzcjy
|
168a31eb00
|
Support prefill max requests limitation (#14993)
|
2025-12-14 10:52:10 +08:00 |
|
fzyzcjy
|
8cc77261ec
|
Super tiny remove unused argument (#14966)
|
2025-12-13 09:41:01 +08:00 |
|
roikoren755
|
2ce121a1c3
|
Enable RadixCache for Mamba2 models (#13584)
|
2025-12-05 18:23:58 +08:00 |
|
  ![github-actions[bot] <github-actions[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)
|
64092c8b55
|
[Auto Sync] Rename is_hybrid to is_hybrid_swa (#14252)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: Hanming Lu <hanming@x.ai>
|
2025-12-01 23:24:24 -08:00 |
|
PiteXChen
|
dc7bdc7329
|
bugfix[schedule]: Excessive preemption occurs when preempting running requests to schedule new prefill requests. (#12494)
Signed-off-by: CLFutureX <chenyongqyl@163.com>
|
2025-11-30 22:29:26 +08:00 |
|
 Michelle Wuandronnie_zheng
|
262c3c1fde
|
[Ascend] Support enable-mixed-chunk in non-MLA scenarios (#12491)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2025-11-27 00:00:21 +08:00 |
|
  ![github-actions[bot] <github-actions[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)
|
e83bd1fadc
|
[Auto Sync] Update schedule_batch.py, schedule_policy.py, b... (20251122) (#13763)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: Hanming Lu <hanming@x.ai>
|
2025-11-24 14:33:31 -08:00 |
|
Liangsheng Yin
|
b2f7b08c49
|
Refactor cache init logic (#13800)
|
2025-11-24 11:41:46 +08:00 |
|
liuhuijiayou
|
dbf22152d6
|
Fix bug: Incorrect variable used in rem_total_token_offset calculatio… (#13201)
|
2025-11-24 11:04:34 +08:00 |
|
Liangsheng Yin
|
ac43822634
|
Refactor eagle bigram key matching (#13714)
|
2025-11-22 20:40:42 +08:00 |
|
lixiaolx
|
d368c7451a
|
(1/n)support context parallel with deepseekv3.2-DSA (#12065)
|
2025-11-16 20:12:25 -08:00 |
|
 cctryandLiangsheng Yin
|
b0b4f71679
|
[Fix] memory leak by overlap + retract (#11981)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-10-23 22:59:23 +08:00 |
|
  
|
a55cf5304a
|
[Feature] Support mamba radix cache v0 (#11214)
Co-authored-by: hanming-lu <hanming@x.ai>
Co-authored-by: hzh0425 <hzh0425@apache.org>
Co-authored-by: thalahors <ericalcaide1@gmail.com>
|
2025-10-12 20:57:15 -07:00 |
|
cctry
|
f3764c26a3
|
Clean match_prefix and prepare_for_extend for mem cache V2 (#11200)
|
2025-10-07 17:54:18 -07:00 |
|
Ke Bao
|
31b49c0b51
|
EAGLE cache fix for HiCache (#11215)
|
2025-10-04 16:53:53 -07:00 |
|
Lianmin Zheng
|
2d62af6be5
|
Fix metrics and request tracing (TimeStats) (#11123)
|
2025-10-01 13:03:07 -07:00 |
|
Xinyuan Tong
|
12d6cf18f0
|
Refactors radix cache for extra key support (#10317)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-09-22 02:16:16 +08:00 |
|
 
|
8ecef73f12
|
[1/2] Support deterministic inference with flashinfer attention backend (#10645)
Co-authored-by: hebiao064 <hebiaobuaa@gmail.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
|
2025-09-19 23:34:29 -07:00 |
|
harrisonlimh
|
14fdd52740
|
feat: add priority based scheduling with priority based request acceptance and preemption (#8746)
|
2025-09-16 17:10:10 -07:00 |
|
 ssshinigamiandMaksim
|
5dd8c6444b
|
[Bug fix] Fix Gemma 2 and fix Gemma 3 multimodal with bs > 1 on NPU (#9871)
Co-authored-by: Maksim <makcum888e@mail.ru>
|
2025-09-08 01:19:40 -07:00 |
|
pansicheng
|
c67569491c
|
Ensure chunked request extension length respects both rem_chunk_tokens and rem_total_tokens limits (#10003)
|
2025-09-04 04:15:26 -07:00 |
|
Liangsheng Yin
|
f9afa7dceb
|
Fix docs for clip max new tokens (#9082)
|
2025-08-11 13:15:21 -07:00 |
|
Yusong Gao
|
4bec99ecd0
|
Fix: resolve prefill of retracted request out-of-memory issue when ignore_eos is enabled (#7434)
|
2025-08-02 14:43:45 +08:00 |
|
 Hanming LuandYing Sheng
|
9379da77de
|
SWA Prefix Cache (#7367)
Co-authored-by: Ying Sheng <sqy1415@gmail.com>
|
2025-07-13 12:31:07 -07:00 |
|
Liangsheng Yin
|
25549433e8
|
Fix prefill OOM due to wrong token calculation when page > 1 (#7397)
|
2025-06-24 02:12:29 +08:00 |
|
Liangsheng Yin
|
05c9bc8956
|
[minor] simplify the TokenToKVPoolAllocator (#7414)
|
2025-06-22 12:37:18 +08:00 |
|
 DarkSharpnessandZhiqiang Xie
|
47367b768d
|
[Refactor] Clean up radix cache related API (#7303)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2025-06-20 00:58:48 +08:00 |
|