fzyzcjy
|
d5431ff894
|
Tiny add stuck simulation (#15613)
|
2025-12-22 17:00:18 +08:00 |
|
fzyzcjy
|
454a2544f2
|
Support soft watchdog for tokenizer/detokenizer/dp-controller processes (#15607)
|
2025-12-22 16:59:16 +08:00 |
|
Hexq0210
|
cb30d056e3
|
add decode round robin policy (#15164)
|
2025-12-22 14:48:52 +08:00 |
|
Chenxi Li
|
0bf95e6d69
|
Fix type mismatch in LoRA batch validation causing assertion failures (#15427)
|
2025-12-21 15:54:41 -08:00 |
|
 
|
bed301a5ac
|
[Feature] Enable return routed experts (#12162)
Co-authored-by: yizhang2077 <1109276519@qq.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-12-21 15:16:43 +08:00 |
|
Xinyuan Tong
|
0a346d3bd9
|
feat: Add limit-mm-data-per-request argument to server arguments (#15418)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-12-20 10:44:47 -08:00 |
|
 ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)     
|
1f1f05a85e
|
vlm: refactor engine vlm params and support processor output as input (#14091)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: zhaochenyang20 <zhaochenyang20@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: BenYao21 <cyao22@asu.edu>
Co-authored-by: minleminzui <minleminzui@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
|
2025-12-20 18:31:24 +08:00 |
|
Liangsheng Yin
|
b9ebf0ed63
|
Clean hidden_states_before_norm (#15485)
|
2025-12-20 09:38:46 +08:00 |
|
Liangsheng Yin
|
933cef16cc
|
Tiny fix mimo model conflicts with main (#15483)
|
2025-12-19 23:20:59 +08:00 |
|
shuwenn
|
5a0ad7310e
|
fix: update model name after weights update (#15416)
|
2025-12-19 21:53:14 +08:00 |
|
 ![github-actions[bot] <github-actions[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)
|
92e6b3c30e
|
[Auto Sync] Update scheduler_runtime_checker_mixin.py (20251219) (#15437)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
|
2025-12-19 01:46:03 -08:00 |
|
 shuwennandLiangsheng Yin
|
fb17845723
|
fix: unreachable error check in retraction (#15433)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-12-19 16:33:34 +08:00 |
|
Lianmin Zheng
|
f228b662a7
|
Update readme (#15425)
|
2025-12-18 23:06:00 -08:00 |
|
+6        
|
160a06cab2
|
[Feature] Xiaomi MiMo-V2-Flash day0 support (#15207)
Co-authored-by: 谢学扬 <xiexueyang@xiaomi.com>
Co-authored-by: tz <tangzhen3@xiaomi.com>
Co-authored-by: 李家乐 <lijiale10@xiaomi.com>
Co-authored-by: 张晨 <zhangchen50@xiaomi.com>
Co-authored-by: Shaohui Liu <liushaohui3@xiaomi.com>
Co-authored-by: 王晨 <wangchen77@xiaomi.com>
Co-authored-by: jiangzihan <jiangzihan@xiaomi.com>
Co-authored-by: xiexueyang <xyxie_wangyi@163.com>
Co-authored-by: Linghao Zhang <zhanglinghao@xiaomi.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: JoyFuture <35593546+JoyFuture@users.noreply.github.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
Co-authored-by: root <root@bj9-ml-g8h20e-k8s-slave106-20251106.alicn.idc.xiaomi.com>
|
2025-12-19 11:40:07 +08:00 |
|
Zehuan Li
|
f6c9db4bc4
|
[DLLM] Fix dLLM regression (#15371)
|
2025-12-19 10:55:46 +08:00 |
|
Zehuan Li
|
b2803ff207
|
[DLLM] Add CI for diffusion LLMs (#14723)
|
2025-12-19 08:54:03 +08:00 |
|
Feng Su
|
29e8f7f9e5
|
multimodal: precompute hash for MultimodalDataItem (#14354)
Signed-off-by: Feng Su <sufeng@linux.alibaba.com>
Signed-off-by: Junjie Mao <junjie.mao@linux.alibaba.com>
|
2025-12-18 15:27:59 -08:00 |
|
fzyzcjy
|
88a405cc10
|
Support EPLB balancedness prometheus metric without GPU->CPU synchronize (#15401)
|
2025-12-18 22:24:23 +08:00 |
|
fzyzcjy
|
602fe3b296
|
Super tiny add moe_ep_rank to prometheus labels (#15407)
|
2025-12-18 22:23:20 +08:00 |
|
fzyzcjy
|
ad9616f13a
|
Tiny extract ModelRunnerOutput (#15400)
|
2025-12-18 22:18:45 +08:00 |
|
fzyzcjy
|
c5f4e20f2f
|
Support GPU execution time breakdown by forward mode metrics (#15396)
|
2025-12-18 22:18:17 +08:00 |
|
Lianmin Zheng
|
d1f0063262
|
Clean up __init__ function of the scheduler and event loop for PD (#15298)
|
2025-12-18 01:35:14 -08:00 |
|
 Xuchun ShangandShangming Cai
|
fea2d5211d
|
[bug fix][pp] fix inconsistent latency between tp (#15379)
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2025-12-18 16:59:56 +08:00 |
|
Shangming Cai
|
ee1ca51de8
|
[PP] Fix dynamic chunking strategy for PP (#15372)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2025-12-18 14:24:55 +08:00 |
|
Scott Lee
|
169a75dfb8
|
Add request-level timestamp for when prefill finishes (#14860)
|
2025-12-17 13:34:46 -08:00 |
|
Liangsheng Yin
|
0c00220795
|
tiny unify environ usage (#15335)
|
2025-12-17 23:31:43 +08:00 |
|
Shangming Cai
|
eeb2b9b259
|
[PP] Minor code cleanup for Pipeline Parallelism (#15329)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2025-12-17 23:31:25 +08:00 |
|
fzyzcjy
|
3e690cce53
|
Add realtime token counter metrics (#15198)
|
2025-12-17 21:20:24 +08:00 |
|
fzyzcjy
|
b12b40dea4
|
Add cuda_graph_forward_passes_total and num_retracted_reqs_total (#15189)
|
2025-12-17 21:19:42 +08:00 |
|
 
|
45a959d3e9
|
[PP] Add pp support for Qwen3-VL (#12333)
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Signed-off-by: Kun(llfl) <i@imux.top>
Signed-off-by: Kun(llfl) <llfl@linux.alibaba.com>
Co-authored-by: kun-llfl <i@imux.top>
Co-authored-by: Kun(llfl) <llfl@linux.alibaba.com>
|
2025-12-17 16:03:58 +08:00 |
|
luqitao
|
0129c911e0
|
fix(function_call): fallback to decode when batch decode options differ (#15155)
Signed-off-by: luqitao <luqt@awcloud.com>
|
2025-12-16 20:21:08 -08:00 |
|
Lianmin Zheng
|
9d64a7b24f
|
Minor style fixes to the scheduler.py (#15218)
|
2025-12-16 17:09:44 -08:00 |
|
amysaq2023
|
ccc8f3b266
|
support non disturbing remote instance weight loader v2 (#14997)
Signed-off-by: Anqi Shen <amy.saq@antgroup.com>
|
2025-12-16 14:39:56 -08:00 |
|
Liangsheng Yin
|
ecb401ed42
|
Enhance runtime memory check in CI (#15192)
|
2025-12-16 21:40:38 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
3912ee4991
|
[VLM] feat: support chunked vit attention (#14907)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-12-15 12:11:02 +08:00 |
|
Hanming Lu
|
e61dabf5e4
|
[Qwen3-next] support mamba radix cache for overlap scheduler (#14792)
|
2025-12-14 18:54:16 -08:00 |
|
    
|
9acb21ae27
|
feat: support EPD disaggregation (#12263)
Co-authored-by: liusy58 <liusy58@linux.alibaba.com>
Co-authored-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: Nicholas <45984215+liusy58@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
|
2025-12-14 22:30:08 +08:00 |
|
liupeng374
|
f50af32d8b
|
[scheduler] remove scheduler allgather for best throughout (#14294)
|
2025-12-14 16:37:49 +08:00 |
|
fzyzcjy
|
168a31eb00
|
Support prefill max requests limitation (#14993)
|
2025-12-14 10:52:10 +08:00 |
|
fzyzcjy
|
fdc93b019c
|
Add sglang:decode_sum_seq_lens metric (#15066)
|
2025-12-14 08:42:45 +08:00 |
|
liupeng374
|
3134d2b2a7
|
[scheduler] enhance scheduler in dp_attention mixed case with spec (#14201)
|
2025-12-14 02:52:26 +08:00 |
|
Liangsheng Yin
|
ed52d01b0b
|
Fix spec info's filter when reqs are finished right after prefill (#14742)
|
2025-12-14 00:32:54 +08:00 |
|
Liangsheng Yin
|
01e3b3f3a3
|
Fix decode OOM caused by retraction (#14939)
|
2025-12-13 12:59:17 +08:00 |
|
fzyzcjy
|
313f59ad80
|
Add soft watchdogs to debug soft hangs (#15023)
|
2025-12-13 10:41:35 +08:00 |
|
fzyzcjy
|
487cf81a68
|
Tiny extract SchedulerWatchdog (#15021)
|
2025-12-13 10:40:32 +08:00 |
|
fzyzcjy
|
8cc77261ec
|
Super tiny remove unused argument (#14966)
|
2025-12-13 09:41:01 +08:00 |
|
Chenxi Li
|
9b9d21312a
|
Feature/Fix multi lora scheduler blocking issue and evict LoRA None lastly (#14795)
|
2025-12-12 17:13:05 -08:00 |
|
 Yineng Zhangandfzyzcjy
|
4b7b5af36a
|
Revert several PRs (#14958)
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
|
2025-12-12 11:25:12 -08:00 |
|
 Kevin LiandShangming Cai
|
8fa8d9d7e8
|
[PD] Add decode PP event loop for PD disaggregation (#14945)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2025-12-12 20:15:00 +08:00 |
|
    
|
c01b2ee094
|
[PP] Refactor PP to async mode (#11852)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: bluecoffee8 <jasperli2002@gmail.com>
Co-authored-by: zhangxiaolei123456 <zhangxiaolei.666@bytedance.com>
Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com>
|
2025-12-12 12:54:16 +08:00 |
|