Cherry_ming
|
1808df48fe
|
[NPU]add nightly-test-npu (#14143)
|
2025-12-05 00:43:35 +08:00 |
|
Liangsheng Yin
|
441420e149
|
Add mooncake transfer_engine_bench into maunal test (#14429)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-04 22:35:00 +08:00 |
|
jianan-gu
|
70d2587324
|
[CPU] Optimize small oc GEMM for Qwen3-next on CPU (#12446)
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com>
|
2025-12-04 00:38:47 -08:00 |
|
Even Zhou
|
894c0dc57c
|
[NPU][1/N] NPU basic functions refactor and new modelslim quant type (#13359)
|
2025-12-04 16:15:31 +08:00 |
|
Ma Mingfei
|
f90b400431
|
[CPU] add support for mamba causal conv1d for qwen3-next (#12309)
|
2025-12-04 13:41:42 +08:00 |
|
sunxxuns
|
5bbd83a2c8
|
ci: Migrate AMD workflows to new MI325 runners; temporarily disabled failed CI's to be added back (#14226)
|
2025-12-03 11:33:27 -08:00 |
|
Lianmin Zheng
|
46d7b35ec7
|
Move custom_ops under layers; move _custom_ops.py → custom_all_reduce_ops.py (#14326)
|
2025-12-03 10:33:37 -08:00 |
|
blzheng
|
974c562a25
|
[CPU] add fused_qkvzba_split_reshape_cat kernel for Qwen3-next (#12330)
|
2025-12-03 23:46:08 +08:00 |
|
Liangsheng Yin
|
24903b88ba
|
Tiny adjust CI testcases (#14362)
|
2025-12-03 21:19:51 +08:00 |
|
ZhengdQin
|
d122e32467
|
[NPU] bug fix: w_vc need contiguous for NPU batch_matmul_transpose ops (#13980)
|
2025-12-03 19:35:18 +08:00 |
|
Cheng Wan
|
96cc10834a
|
[CI] update estimated elapsed time of some unittests (#14347)
|
2025-12-03 01:21:40 -08:00 |
|
Yuhao Yao
|
77512ae0d7
|
[bugfix] Fix prefill tbo disabled when --deepep-mode=auto (#14333)
Co-authored-by: Cheng Wan <wan4ch@gmail.com>
|
2025-12-03 01:20:33 -08:00 |
|
Xuan Liao
|
c233e9d7a9
|
[CPU] Support chunk_gated_delta_rule kernel for Qwen3-Next (#12441)
|
2025-12-03 17:03:48 +08:00 |
|
Johnsonms
|
043f13171f
|
[Performance] Optimize NSA Indexer K/S Buffer Access with Fused Triton Kernels (#13812)
Co-authored-by: Johnsonms <johnson@together.ai>
|
2025-12-02 18:53:06 -08:00 |
|
sglang-bot
|
7ae368efde
|
chore: bump SGLang version to 0.5.6 (#14316)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2025-12-02 17:17:13 -08:00 |
|
Eva20150932-atlascloud
|
7c38eca1e4
|
feat: DeepSeek new v3.2 encoding (#14249)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-12-02 11:41:05 -08:00 |
|
Lianmin Zheng
|
64092c8b55
|
[Auto Sync] Rename is_hybrid to is_hybrid_swa (#14252)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: Hanming Lu <hanming@x.ai>
|
2025-12-01 23:24:24 -08:00 |
|
Yuan Luo
|
26aebf83d3
|
[VLM] Support Piecewise CUDA Graph for Qwen3-Omni-MOE (#14222)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-12-02 10:12:10 +08:00 |
|
Baizhou Zhang
|
eb5008846a
|
[CI] Fix test_deepep_large.py (#14247)
|
2025-12-01 15:18:48 -08:00 |
|
liupeng374
|
2e8f54e61e
|
[spec-overlap] bugfix for pd disaggregation and npu (#14088)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
|
2025-12-01 22:58:20 +08:00 |
|
YAMY
|
decb48965d
|
[DeepSeekV3.2] Enable pure TP & Partial DP Attention (#13646)
|
2025-11-30 15:59:23 -08:00 |
|
Liangsheng Yin
|
0a9d64530d
|
Support grammar + spec + reasoning (#14163)
|
2025-11-30 21:19:57 +08:00 |
|
fzyzcjy
|
36b729c2b8
|
Implement profiler v2 and fix stage mixture bug (#14148)
|
2025-11-30 16:59:52 +08:00 |
|
Kangyan-Zhou
|
1d3d8b3418
|
Fix Minimax M2 loading issue (#13956)
|
2025-11-29 17:07:19 -05:00 |
|
Lianmin Zheng
|
155a9e7237
|
Fix condition for streaming output_ids in tokenizer manager (#13759)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Chang Su <chang.s.su\n@oracle.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2025-11-29 13:56:15 -08:00 |
|
Shangming Cai
|
c6d34a0688
|
Move piecewise cuda graph test to manual dir to fix CI (#14121)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2025-11-29 00:52:01 -08:00 |
|
Cheng Wan
|
0fe74af563
|
Remove incorrect deep_gemm assertions from server_args.py (#14113)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
|
2025-11-28 20:25:39 -08:00 |
|
fjybiocs
|
143b57b805
|
enable piecewise cuda graph for prefill server (#13377)
Co-authored-by: serverance.fu <serverance.fu@temu.com>
|
2025-11-29 12:09:26 +08:00 |
|
Yan Ru Pei
|
f446b51c41
|
fix: malformed KV events for NVIDIA Dynamo (#13488)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
|
2025-11-28 14:55:20 -08:00 |
|
Kangyan-Zhou
|
a102a0507a
|
Disable Deepep 2 GPU tests (#14111)
|
2025-11-28 13:56:34 -08:00 |
|
Lzhang-hub
|
ea1e9f6b3c
|
feat: support qwen3_vl vision model dp (#13724)
|
2025-11-28 17:29:07 +08:00 |
|
Raayan Dhar
|
91d249cd9b
|
fix: small changes to enable test_mrope.py (#14082)
Signed-off-by: Raayan Dhar raayan.dhar@gmail.com <raayan.dhar@gmail.com>
|
2025-11-27 22:58:17 -08:00 |
|
alisonshao
|
bce40fa217
|
Fix utils import issue for nightly tests (#13944)
|
2025-11-27 15:24:27 -08:00 |
|
Kangyan-Zhou
|
4c9f7c97d3
|
Temporarily disabled test (#14069)
|
2025-11-27 14:58:46 -08:00 |
|
Yixin Dong
|
6350042696
|
feat: Naive support Spec V2 + Constrained Decoding (#13425)
Signed-off-by: Ubospica <ubospica@gmail.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-11-27 20:31:46 +08:00 |
|
alisonshao
|
d941a3befa
|
Fix nightly test failure: NSA indexer dtype (#14017)
|
2025-11-26 21:27:51 -05:00 |
|
Sam
|
91e8dc371a
|
[Feat][NVFP4] Enable NVFP4 MoE for Qwen series models (eg. Qwen3-Next) #13761 (#13761)
Co-authored-by: Kaixi Hou <kaixih@nvidia.com>
|
2025-11-26 17:53:45 -07:00 |
|
ShawnY112358
|
5155016b56
|
[feat] update bucketed weights from distributed (#13824)
Co-authored-by: Stefan He <hebiaobuaa@gmail.com>
|
2025-11-26 15:30:45 -08:00 |
|
Netanel Haber
|
082b54c689
|
Support nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16 (and nvidia/C-RADIOv2-H) (#12277)
|
2025-11-26 16:28:52 -07:00 |
|
alisonshao
|
a8ef4d1804
|
Add nightly test support to unified run_suite.py (#13941)
|
2025-11-26 15:16:12 -08:00 |
|
alisonshao
|
5b7da0f58e
|
Temporarily disable test_update_weights_from_disk.py in CI (#14021)
|
2025-11-26 13:56:28 -08:00 |
|
Douglas Yang
|
697a77bf73
|
Add stress test workflow (#13937)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2025-11-26 13:24:39 -08:00 |
|
Liangsheng Yin
|
6c190cbda0
|
Rename: --hooks to --forward-hooks (#13994)
|
2025-11-26 22:26:28 +08:00 |
|
StonyPort
|
540d6fee20
|
Support piecewise CUDA graph for embedding models (#13852)
Co-authored-by: qiuxuan.lzw <qiuxuan.lzw@alibaba-inc.com>
|
2025-11-26 15:29:46 +08:00 |
|
ShawnY112358
|
007c3e234c
|
[feat] support in-flight weight update (#10071)
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
|
2025-11-25 22:03:13 -08:00 |
|
Yubo Wang
|
18fb51583f
|
Support FlashAttention3 page_size > 1 and topk > 1 case with paged attn and spec decode (#7725)
|
2025-11-26 11:44:41 +08:00 |
|
Yuan Luo
|
ca5c8b16f6
|
[VLM] Support InternVL Vision Encoder Data Parallelism (#13925)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-11-26 11:43:05 +08:00 |
|
Fan Yin
|
36b1bcd242
|
[chore] update torch version to 2.9 (#12969)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-11-25 14:47:34 -08:00 |
|
Lianmin Zheng
|
1ab6ce0e62
|
[Auto Sync] Improve profilers and simplify bench_one_batch_server.py (#13866)
|
2025-11-25 12:13:31 -08:00 |
|
Liangsheng Yin
|
3421d049aa
|
[CI] rename: per_commit -> registered (#13928)
|
2025-11-25 23:05:06 +08:00 |
|