Commit Graph
5274 Commits
Author SHA1 Message Date
398b81f78c Support GlmMoeDsaForCausalLM (#18521)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Signed-off-by: BBuf <1182563586@qq.com>
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
Co-authored-by: BBuf <1182563586@qq.com>
2026-02-10 15:20:10 +08:00
e8a2c13380 Deepseekv32 compatibility with transformers v5 (#18297)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-02-10 14:50:40 +08:00
siyu 0b15f19927 [EPD] Add notification mechanism to fix server hang and add timeout env var (#18229) 2026-02-10 11:52:54 +08:00
Qiaolin Yu 4a1b50bb2d Fix idle batch predict dtype in spec v2 (#18379) 2026-02-10 10:29:13 +08:00
Kartik Ramesh 26a006e47f Add cache_config_info metric. (#17273) 2026-02-09 16:09:09 -08:00
Lianmin Zhenggithub-actions[bot] <github-actions[bot]@users.noreply.github.com>Hanming Lugemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
b027c5aca6 [Auto Sync] Update cache_init_params.py (20260209) (#18502)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-09 14:49:41 -08:00
Lianmin Zhengandgithub-actions[bot] <github-actions[bot]@users.noreply.github.com> ce95f203b0 [Auto Sync] Update logits_processor.py (20260209) (#18503)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-02-09 14:33:02 -08:00
ishandhanani 01e3f4682e feat(kv-events): Add medium field to KV event types for storage tier tracking (#18205) 2026-02-09 12:39:15 -08:00
Zheng Liand瑀澈 27c447653d model: support Qwen3.5 (#18489)
Co-authored-by: 瑀澈 <yuche.lz@alibaba-inc.com>
2026-02-10 00:27:59 +08:00
Kurt Shuster 006da22268 Pass quantize_config to _initialize_model (#18273) 2026-02-09 23:34:42 +08:00
brimon ddbcfbaaab feature: support bidirectional attention for Gemma-3 (#10707) 2026-02-09 23:17:45 +08:00
Baizhou Zhang 615a02dcd4 Revert "optimize get_topk_ragged by fusing get k and k_scale triton kernel" (#18471) 2026-02-09 16:37:19 +08:00
Liangsheng Yin 875ad6cf35 Tiny rename for spec related fileds. (#18468) 2026-02-09 00:10:39 -08:00
LHXuuuPeng Zhanggemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
107958a489 Make compressed-tensors MoEs support ignored layers (#17828)
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-09 14:37:33 +08:00
Junlin ZhouandTiwei Bie 14652243bd [DLLM] Add JointThreshold algorithm for joint M2T and T2T decoding (#18171)
Signed-off-by: Junlin Zhou <zhoujunlin.zjl@antgroup.com>
Co-authored-by: Tiwei Bie <tiwei.btw@antgroup.com>
2026-02-09 14:20:45 +08:00
3f3c201243 [AMD] Update aiter to v0.1.10.post2 (#18423)
Co-authored-by: kkHuang-amd <wunhuang@amd.com>
Co-authored-by: YC Tseng <yctseng@amd.com>
2026-02-08 22:08:24 -08:00
Yingchun LaiandLiangsheng Yin a1189068fa fix: fix the wrong return value type of draft model runner (#18105)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2026-02-08 20:51:35 -08:00
Zheng Wengang 68e31a3485 [BugFix][PD]Fix metadata_buffer_index leak when aborted in PD (#17483) 2026-02-09 11:34:29 +08:00
Shangming Cai bffd765417 Refactoring Mooncake TE as a shared distributed component (#17810)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-02-09 10:53:11 +08:00
Yi Zhong bf89cc3803 [ModelOPT] Support Qwen 3 Next Coder NVFP4 (#18224)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
2026-02-08 22:29:07 +00:00
Mohammad Miadh Angkad 071bf2ce09 [Kimi-K2.5] Fix missing quant_config in KimiK25 (#18440) 2026-02-08 12:02:45 -08:00
Piotr Mazurek 656a3d742e Add tensor parallelism support to LFM2 ShortConv layers (#17777) 2026-02-09 00:52:47 +08:00
debo3 031a652b93 Fix TRT-LLM MLA backend applying k_scale to BF16 KV cache in BMM1 (#18396) 2026-02-08 23:11:16 +08:00
Yi Zhong ca36d88fa6 [ModelOpt] Fix broken Qwen3-235B-A22B-Instruct-2507-NVFP4 launch (#18189)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
2026-02-08 14:35:28 +00:00
Zack Yu d71ccd8860 fix: sync server_args.kv_cache_dtype when detecting FP8 KV cache (#18394) 2026-02-08 14:10:59 +08:00
DarkSharpness 8e2e835c2f [Fix] Fix backend selection after flashinfer version update (#18364) 2026-02-08 11:20:41 +08:00
Mohammad Miadh Angkad 7b83659310 fix: fix NVFP4 Kimi-K2.5 weight mapping and exclude list (#18370) 2026-02-08 10:23:48 +08:00
Mohammad Miadh Angkad fddef76619 [Doc] Fix outdated --fp4-gemm-backend documentation (#18350) 2026-02-07 20:42:47 +08:00
hlu1 4637970dfb [Qwen3Next] Optimize fused_sigmoid_gating_delta_rule_update_kernel (#18271) 2026-02-07 11:59:42 +08:00
Neal Vaidyaandishandhanani f1ff697494 add hybrid model PD to NIXL connector (#16229)
Signed-off-by: Neal Vaidya <nealv@nvidia.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
2026-02-06 15:05:36 -08:00
Alison Shao d0c39bc219 Fix cross-container HF download race condition in CI (#18328) 2026-02-05 21:01:41 -08:00
Linyu Wu aa390d2762 [Kernel] Migrate GPTQ-Marlin GEMM kernel to JIT (#18067) 2026-02-06 08:31:42 +08:00
6a4b81e2d9 Refactor(qwen3-vl) optimize position encoding interpolation (#16781)
Signed-off-by: chenzhenyang <andy271828@163.com>
Signed-off-by: chenzhenyang <chenzhenyang@moonshot.cn>
Co-authored-by: chenzhenyang <chenzhenyang@moonshot.cn>
Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>
2026-02-05 10:26:35 -08:00
ovidiusm 498d8d0680 NixlKVManager optimizations (#17654)
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com>
2026-02-06 00:25:23 +08:00
Glen Liu 3f32a5831d throw error if got adapter with added_tokens (#18046) 2026-02-05 23:55:43 +08:00
pansicheng 2eb4359ada [Kernel] Add JIT apply_rope_with_cos_sin_cache_inplace (#18155) 2026-02-05 21:49:37 +08:00
Shangming Cai afae4c7178 [PD] Minor code cleanup for mooncake backend (#18279)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-02-05 17:38:09 +08:00
zhangheng 079fc8f3c5 [piecewise graph]: support MiniMax-M2 (#18217) 2026-02-04 23:24:38 -08:00
danielafrimi 3f1df322f9 [FIX] Always support TP > 4 for FP4 Gemm (#17300) 2026-02-05 15:10:26 +08:00
Meng, Hengyu 368936a62b [XPU] Integrate MoE and minor improvements in XPU attention backend (#13561) 2026-02-04 23:09:59 -08:00
Ch3ngY1andShangming Cai f730c18679 [PD] improve kv offset calculation for MHA model with different tp size (#18163)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-02-05 10:43:23 +08:00
yinghui 599c5f4922 fix kimi k2.5's moe gemm config init (#18064) 2026-02-04 16:59:01 -08:00
Mohammad Miadh Angkad efbf39583e Add MoE fused config for Qwen3-Coder-Next-FP8 on H100 TP=2 (#18195) 2026-02-04 13:36:35 -08:00
RunningLeonandKe Bao 3e7ecb78a6 model: support interns1-pro (#18145)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
2026-02-05 00:22:44 +08:00
RunningLeon a6f53cc5e3 entrypoint: support passing spaces_between_special_tokens per request (#17939) 2026-02-04 22:18:36 +08:00
BingjiaWangandabing 760ae933bb optimize get_topk_ragged by fusing get k and k_scale triton kernel (#16043)
Co-authored-by: abing <wangbingjia.wbj@alibaba-inc.com>
2026-02-04 19:59:41 +08:00
Nicolas Castet 315306d8a9 Make sure we always disable symm memory without dp padding (#18129) 2026-02-04 19:58:28 +08:00
Jincong Chen a72f4f839c Tiny fix for fp8 moe backend flashinfer_trtllm naming (#18243) 2026-02-04 19:58:04 +08:00
Cheng Wan 84c09913eb Moving _alloc_extend_naive out of npu allocator (#18200) 2026-02-04 02:09:55 -08:00
zhangheng be557cbc5f [RadixTree][5/N Refactor]: Introduce pre and post-processing methods for key matching (#18147) 2026-02-04 17:10:46 +08:00