Baizhou Zhang
|
6799847ebf
|
[CI]Unblock and split spec v2+dp test (#14551)
|
2025-12-07 17:39:25 -08:00 |
|
Baizhou Zhang
|
673c11ba73
|
[Minor] Temporarily skipping deepep large mtp test (#14586)
|
2025-12-07 13:59:16 -08:00 |
|
b8zhong
|
3b47973af8
|
[CI] Tiny speed up VLM CI (#14517)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
|
2025-12-07 13:30:41 -08:00 |
|
khalilzhk
|
948b6acee8
|
[BugFix] fix prefixcache performance and accuracy on ascend (#13573)
|
2025-12-08 02:16:20 +08:00 |
|
Liangsheng Yin
|
125e17efd5
|
Add small model test for spec v2 + dp + trtllm_mla (#14576)
|
2025-12-07 23:55:00 +08:00 |
|
Hanming Lu
|
e592ee6545
|
[Qwen3-next] remove heuristics and add radix cache kl test (#14520)
|
2025-12-06 12:11:40 -08:00 |
|
Baizhou Zhang
|
bc388471d2
|
[1/n] Fix hanging during DeepGemm Warmup (#14493)
|
2025-12-06 10:44:02 -08:00 |
|
Baizhou Zhang
|
42fcf5438f
|
Revert "tiny remove deprecated endpoint call" (#14533)
|
2025-12-05 23:48:54 -08:00 |
|
blzheng
|
d257bf87b9
|
[CPU] add mamba fla kernels for Qwen3-next (#12324)
|
2025-12-06 14:16:23 +08:00 |
|
fzyzcjy
|
3d1b591aa1
|
Tiny use trtllm_mha as default when possible (#14291)
|
2025-12-05 14:26:03 -08:00 |
|
Alison Shao
|
16e8463a90
|
Add Mistral Large 3 basic test to PR CI (#14460)
|
2025-12-05 10:38:52 -08:00 |
|
b8zhong
|
ec7b2c16d9
|
tiny remove deprecated endpoint call (#13607)
|
2025-12-05 09:54:49 -08:00 |
|
roikoren755
|
889b46ea50
|
[Spec] Mamba2 support in target models (#13434)
|
2025-12-06 00:50:46 +08:00 |
|
Xinyuan Tong
|
6d37e70883
|
ministral3 (#14251)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Yueming Yuan <yy28@illinois.edu>
|
2025-12-04 14:31:26 -08:00 |
|
jianan-gu
|
70d2587324
|
[CPU] Optimize small oc GEMM for Qwen3-next on CPU (#12446)
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com>
|
2025-12-04 00:38:47 -08:00 |
|
Even Zhou
|
894c0dc57c
|
[NPU][1/N] NPU basic functions refactor and new modelslim quant type (#13359)
|
2025-12-04 16:15:31 +08:00 |
|
Ma Mingfei
|
f90b400431
|
[CPU] add support for mamba causal conv1d for qwen3-next (#12309)
|
2025-12-04 13:41:42 +08:00 |
|
sunxxuns
|
5bbd83a2c8
|
ci: Migrate AMD workflows to new MI325 runners; temporarily disabled failed CI's to be added back (#14226)
|
2025-12-03 11:33:27 -08:00 |
|
blzheng
|
974c562a25
|
[CPU] add fused_qkvzba_split_reshape_cat kernel for Qwen3-next (#12330)
|
2025-12-03 23:46:08 +08:00 |
|
Liangsheng Yin
|
24903b88ba
|
Tiny adjust CI testcases (#14362)
|
2025-12-03 21:19:51 +08:00 |
|
ZhengdQin
|
d122e32467
|
[NPU] bug fix: w_vc need contiguous for NPU batch_matmul_transpose ops (#13980)
|
2025-12-03 19:35:18 +08:00 |
|
Cheng Wan
|
96cc10834a
|
[CI] update estimated elapsed time of some unittests (#14347)
|
2025-12-03 01:21:40 -08:00 |
|
Yuhao Yao
|
77512ae0d7
|
[bugfix] Fix prefill tbo disabled when --deepep-mode=auto (#14333)
Co-authored-by: Cheng Wan <wan4ch@gmail.com>
|
2025-12-03 01:20:33 -08:00 |
|
Xuan Liao
|
c233e9d7a9
|
[CPU] Support chunk_gated_delta_rule kernel for Qwen3-Next (#12441)
|
2025-12-03 17:03:48 +08:00 |
|
Eva20150932-atlascloud
|
7c38eca1e4
|
feat: DeepSeek new v3.2 encoding (#14249)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-12-02 11:41:05 -08:00 |
|
Yuan Luo
|
26aebf83d3
|
[VLM] Support Piecewise CUDA Graph for Qwen3-Omni-MOE (#14222)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-12-02 10:12:10 +08:00 |
|
Baizhou Zhang
|
eb5008846a
|
[CI] Fix test_deepep_large.py (#14247)
|
2025-12-01 15:18:48 -08:00 |
|
liupeng374
|
2e8f54e61e
|
[spec-overlap] bugfix for pd disaggregation and npu (#14088)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
|
2025-12-01 22:58:20 +08:00 |
|
Liangsheng Yin
|
0a9d64530d
|
Support grammar + spec + reasoning (#14163)
|
2025-11-30 21:19:57 +08:00 |
|
fzyzcjy
|
36b729c2b8
|
Implement profiler v2 and fix stage mixture bug (#14148)
|
2025-11-30 16:59:52 +08:00 |
|
Lianmin Zheng
|
155a9e7237
|
Fix condition for streaming output_ids in tokenizer manager (#13759)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Chang Su <chang.s.su\n@oracle.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-11-29 13:56:15 -08:00 |
|
Shangming Cai
|
c6d34a0688
|
Move piecewise cuda graph test to manual dir to fix CI (#14121)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2025-11-29 00:52:01 -08:00 |
|
Cheng Wan
|
0fe74af563
|
Remove incorrect deep_gemm assertions from server_args.py (#14113)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
|
2025-11-28 20:25:39 -08:00 |
|
fjybiocs
|
143b57b805
|
enable piecewise cuda graph for prefill server (#13377)
Co-authored-by: serverance.fu <serverance.fu@temu.com>
|
2025-11-29 12:09:26 +08:00 |
|
Yan Ru Pei
|
f446b51c41
|
fix: malformed KV events for NVIDIA Dynamo (#13488)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2025-11-28 14:55:20 -08:00 |
|
Kangyan-Zhou
|
a102a0507a
|
Disable Deepep 2 GPU tests (#14111)
|
2025-11-28 13:56:34 -08:00 |
|
Raayan Dhar
|
91d249cd9b
|
fix: small changes to enable test_mrope.py (#14082)
Signed-off-by: Raayan Dhar raayan.dhar@gmail.com <raayan.dhar@gmail.com>
|
2025-11-27 22:58:17 -08:00 |
|
alisonshao
|
bce40fa217
|
Fix utils import issue for nightly tests (#13944)
|
2025-11-27 15:24:27 -08:00 |
|
Kangyan-Zhou
|
4c9f7c97d3
|
Temporarily disabled test (#14069)
|
2025-11-27 14:58:46 -08:00 |
|
Yixin Dong
|
6350042696
|
feat: Naive support Spec V2 + Constrained Decoding (#13425)
Signed-off-by: Ubospica <ubospica@gmail.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-11-27 20:31:46 +08:00 |
|
ShawnY112358
|
5155016b56
|
[feat] update bucketed weights from distributed (#13824)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
|
2025-11-26 15:30:45 -08:00 |
|
Netanel Haber
|
082b54c689
|
Support nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16 (and nvidia/C-RADIOv2-H) (#12277)
|
2025-11-26 16:28:52 -07:00 |
|
alisonshao
|
5b7da0f58e
|
Temporarily disable test_update_weights_from_disk.py in CI (#14021)
|
2025-11-26 13:56:28 -08:00 |
|
Liangsheng Yin
|
6c190cbda0
|
Rename: --hooks to --forward-hooks (#13994)
|
2025-11-26 22:26:28 +08:00 |
|
StonyPort
|
540d6fee20
|
Support piecewise CUDA graph for embedding models (#13852)
Co-authored-by: qiuxuan.lzw <qiuxuan.lzw@alibaba-inc.com>
|
2025-11-26 15:29:46 +08:00 |
|
ShawnY112358
|
007c3e234c
|
[feat] support in-flight weight update (#10071)
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
|
2025-11-25 22:03:13 -08:00 |
|
Yubo Wang
|
18fb51583f
|
Support FlashAttention3 page_size > 1 and topk > 1 case with paged attn and spec decode (#7725)
|
2025-11-26 11:44:41 +08:00 |
|
Fan Yin
|
36b1bcd242
|
[chore] update torch version to 2.9 (#12969)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-11-25 14:47:34 -08:00 |
|
Chen1022
|
d64bf6c6ce
|
Support piecewise cuda graph for Qwen3-next (#13081)
|
2025-11-25 21:01:27 +08:00 |
|
alisonshao
|
dbab5d50a3
|
Add test_dummy_grok_models.py to not_in_ci section (#13908)
|
2025-11-25 16:51:25 +08:00 |
|