Commit Graph
5553 Commits
Author SHA1 Message Date
fzyzcjy 749736ba95 Tiny fix single-gpu dumper and add tests for dumper (#16285) 2026-01-02 13:30:09 +08:00
Nan JiangandXinyuan Tong 7254986342 [VLM] feat: true on policy for vlm + fsdp (#14636)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-01-01 16:54:39 -08:00
DarkSharpness f6f7af4068 [Refactor] Clean up custom op (#15995) 2026-01-01 21:41:56 +08:00
DarkSharpness a3b1e8ef3d [Feature] add aligned_vector type for JIT kernel (#16162) 2026-01-01 21:40:05 +08:00
fzyzcjy db499e1889 Tiny add filter, support duplications, add visualizations, fix error and robustness for dump comparator (#16262) 2026-01-01 17:47:00 +08:00
fzyzcjy 6cf3a6dd69 Support HTTP control for dumper (#16261) 2026-01-01 17:40:17 +08:00
fzyzcjy 90e24f5c31 Tiny add filter, dump dict, failable save to dumper (#16260) 2026-01-01 17:35:38 +08:00
Mick 21de3e1406 [diffusion] webui: tiny fix loading output image (#16251) 2026-01-01 14:34:49 +08:00
Lianmin Zheng e4c1e441af Fix deprecation warning of sre_parse for python 3.13 (#16247) 2025-12-31 21:25:09 -08:00
e0e5084802 Fix parse args from file(#13911) (#14085)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2026-01-01 11:37:33 +08:00
Yuan Luoandluoyuan.luo 3a42c5e341 [VLM] Adopt jit qk_norm kernel in VLM (#16171)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-01-01 10:10:36 +08:00
Kangyan-Zhou 12b89e51d8 Add P90/99 e2e latency in bench_serving script (#16245) 2025-12-31 15:40:33 -08:00
Feng Su 57d2ba9203 SGLang Tracing: Supports propagating trace headers through sgl.Engine… (#15814)
Signed-off-by: Feng Su <sufeng@linux.alibaba.com>
2025-12-31 15:19:53 -08:00
ishandhanani 0500fea965 fix editable install (#16241) 2025-12-31 14:34:54 -08:00
Cheng Wan 2b461c15b4 Update logprob_start_len handling in scheduler (#16240) 2025-12-31 14:11:24 -08:00
siyu abdf65d4f3 Fix OOM by offloading multimodal features to CPU after embedding (#16018) 2025-12-31 23:02:34 +08:00
Huaixin ChangandLiangsheng Yin c1dfbc777b deprecate prefill-round-robin-balance (#16195)
Signed-off-by: Chang Huaixin (OpenAnolis) <changhuaixin@linux.alibaba.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2025-12-31 22:25:33 +08:00
Muqi Li 2667c857a7 Fix DeepSeekV31's structural tag trigger (#13394) 2025-12-31 21:13:52 +08:00
Kangyan-Zhou fc643ffbc9 Download missing shards in model weights files when not in CI (#16211) 2025-12-31 20:42:34 +08:00
Yuhao YangandMick 4280a18a13 [diffusion] CI: add test for cache-dit (#16204)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-12-31 19:58:10 +08:00
Baizhou Zhang e47afa0237 [DP]Fix sync bubble in adjust_num_token_non_padded_for_attn_tp (#16178) 2025-12-31 17:58:18 +08:00
cen121212 25b48564c3 [NPU][Bugfix] fix Qwen3-VL-30B-A3B-Instruct accuracy loss (#15597) 2025-12-31 15:57:38 +08:00
chhnbandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> 3619ec61b4 [diffusion] feat: support multi-frame image output (#15878)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-31 13:50:02 +08:00
Hudson Xing 8a9ca41fda Move log_prefill_stats_late to correct location in PP mode (#15946) 2025-12-31 12:00:24 +08:00
roikoren755 47a660d5b9 [NemotronH] PP support (#16172)
Signed-off-by: Roi Koren <roik@nvidia.com>
2025-12-31 11:16:15 +08:00
Mickandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> 5bf0d862dd [diffusion] CI: fix generate mode and add cli test (#16174)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-31 09:51:48 +08:00
liupeng374 75b72eb8b2 [cp] assert dsv3.2 cp in pd decode mode (#16156) 2025-12-31 09:24:46 +08:00
bc8b526eda Fix: Handle empty func_name and None values in GLM MoE detectors (#15754)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2025-12-30 14:32:35 -08:00
Izzy Putterman 3dfff6ae3c Eagle: GPT-OSS Eagle v2 support (#14920)
Signed-off-by: Izzy Putterman <iputterman@nvidia.com>
2025-12-30 14:23:07 -08:00
EkiRui ad2c1ee352 doc: mooncake store add dummy client support (#16050)
Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>
2025-12-30 12:37:57 -08:00
Yineng Zhanggithub-actions[bot] <github-actions[bot]@users.noreply.github.com>Scott Lee
00e607111a [Auto Sync] Update request_metrics_exporter.py (20251230) (#16183)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Scott Lee <scottjlee@users.noreply.github.com>
2025-12-30 11:09:41 -08:00
Roger Youngandxuebi d17b9e6392 Fusing RMSNormTP in minimax_m2 (#14416)
Co-authored-by: xuebi <xuebi@minimaxi.com>
2025-12-30 10:22:07 -08:00
Liangsheng Yin ba67e006a7 Refactor speculative algorithm registry. (#16168) 2025-12-31 01:24:22 +08:00
DarkSharpness 45f3ad2f52 [Refactor] Rename CustomOp -> MultiPlatformOp (#16175) 2025-12-31 01:16:32 +08:00
Baizhou Zhang f35b5da521 [CI] Append test variant name to markdown report header in nightly test (#16166) 2025-12-31 00:09:24 +08:00
Xiaoyu Zhangandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> 733a0c1a37 [Diffusion] Zimage opt with qknorm and flashinfer rope (#16161)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-30 23:39:32 +08:00
Xiaoyu Zhangandgithub-actions[bot] <github-actions[bot]@users.noreply.github.com> b369aaa23f [Diffusion] Refine diffusion profling doc (#16163)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2025-12-30 23:38:46 +08:00
Ke Bao b973202526 Split tp model worker init (#16165) 2025-12-30 23:36:03 +08:00
Mufeez Amjad cbff7ad985 dp-attention: add follow_bootstrap_room + auto load-balance; drop decode_round_robin (#16110) 2025-12-30 22:33:06 +08:00
Yuhao YangandMick 39ca57cd28 [diffusion] chore: tiny fix model config (#16159)
Co-authored-by: Mick <mickjagger19@icloud.com>
2025-12-30 22:11:38 +08:00
Mick 3449806727 [diffusion] feat: generalize layer-wise-offload to all supported models (#16150) 2025-12-30 22:06:57 +08:00
Ke Bao b3817fa93b Split model_worker init function (#16160) 2025-12-30 21:39:11 +08:00
Ke Bao 059428bd8a Tiny remove additional args in init_memory_pool (#16158) 2025-12-30 21:38:06 +08:00
jiapingW 5d200dd8d9 [diffusion] bench: distinguish between video generation and image generation in the bench_serving (#16149) 2025-12-30 21:29:18 +08:00
Jumiarandliusy58 664f611e83 Add profiling capture support to the encoder server (#15730)
Signed-off-by: liuanqi <liuanqi6@xiaomi.com>
Co-authored-by: liusy58 <liusy58@linux.alibaba.com>
2025-12-30 19:56:25 +08:00
yh0903Han YuMickgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
49adb37e37 [diffusion] chore: fix ZMQ binding and model loading for FastWan compatibility (#13978)
Co-authored-by: Han Yu <hyu5@dt-login01.delta.ncsa.illinois.edu>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-30 18:09:38 +08:00
Gaoji Liuandliugaoji.lgj 7518dc3532 feat(SpecEagleV2): add standalone_worker_v2 (#12625)
Co-authored-by: liugaoji.lgj <liugaoji.lgj@alibaba-inc.com>
2025-12-30 17:55:04 +08:00
Yuan Luoandluoyuan.luo 94bcc19bce [VLM] Support Video for InternVL3_5 (#15942)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2025-12-30 17:07:49 +08:00
Shangming Cai db3821a9ef [PP] Add a minimum chunk value for PP dynamic chunking (#16140)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2025-12-30 16:35:39 +08:00
Liangsheng Yin c2601f0d21 Fix wrong assigning extend_input_len_per_req with eagle. (#16129) 2025-12-30 15:51:53 +08:00