Alison Shao
|
52c604342c
|
chore: print test list at beginning and end of run_suite.py (#16334)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-01-04 11:52:52 -08:00 |
|
Alison Shao
|
ff0f370f85
|
ci: migrate MoE tests to test/registered/moe/ (#16127)
|
2026-01-04 11:51:12 -08:00 |
|
Alison Shao
|
26f9e20755
|
ci: migrate quantization kernel tests to test/registered/quant/ (#16323)
|
2026-01-04 11:49:29 -08:00 |
|
fzyzcjy
|
2337b1bbb0
|
Refactor and fix prefill delayer (scheduler enhancer) (#16269)
|
2026-01-04 07:15:06 +08:00 |
|
sunxxuns
|
8b869e326c
|
[AMD] feat: add DLLM support for AMD GPUs with LLaDA2 testing (#15560)
|
2026-01-03 10:41:11 +08:00 |
|
Baizhou Zhang
|
f07e76b229
|
Multiple refactors of DeepSeek V32 and context parallel (#16305)
|
2026-01-03 02:21:22 +08:00 |
|
Yineng Zhang
|
417e75a60f
|
fix: use 2-gpu-runner for pp 2 NemotronH ut (#16239)
|
2025-12-31 14:51:33 -08:00 |
|
Bingxu Chen
|
1c360bf753
|
[AMD CI] add testcases to unit-test-backend-1-gpu (#16117)
|
2025-12-30 23:08:45 -08:00 |
|
Vitaly Tuzov
|
1048803c1f
|
Reworked fast_pos_embed_interpolate() using torch (#10959)
|
2025-12-30 14:45:34 +08:00 |
|
Alison Shao
|
f4ec6f8e17
|
ci: migrate remaining spec/eagle tests to test/registered/spec/ (#15800)
|
2025-12-29 14:00:00 -08:00 |
|
Ke Bao
|
faecd37ed4
|
Add Mimo-v2-flash model to ci test (#15887)
|
2025-12-27 14:18:08 +08:00 |
|
Alison Shao
|
72a980c6d5
|
ci: migrate MLA tests to test/registered/mla/ (#15798)
|
2025-12-24 23:29:30 -08:00 |
|
michael-amd
|
e7b09efc0a
|
[AMD] Add AMD Nightly Performance & VLMs Accuracy Tests (#15500)
|
2025-12-23 19:03:27 -08:00 |
|
Liangsheng Yin
|
0d0367e9d0
|
Tiny fix CI (#15696)
|
2025-12-24 01:51:08 +08:00 |
|
Liangsheng Yin
|
bd572360f3
|
Tiny apply gsm8k mixin to ngram test (#15606)
|
2025-12-24 01:30:26 +08:00 |
|
Alison Shao
|
ac42797cf7
|
[CI] Enable retry logic for flaky CI tests (#14983)
|
2025-12-22 22:23:42 -08:00 |
|
Alison Shao
|
883747ced1
|
[CI] Migrate Attention Backend tests to test/registered/attention/ (#15563)
|
2025-12-22 22:17:52 -08:00 |
|
SeanWeiSean
|
7759716786
|
bugfix[schedule]: Refactor sort method and add related UT (#13576)
Co-authored-by: Yuxuan Wei <w1300012920@pku.edu.cn>
Co-authored-by: Yuxuan Wei <w1300012920@gmail.com>
Co-authored-by: Yuxuan Wei🚚 <yuxwei@microsoft.com>
|
2025-12-23 03:13:57 +08:00 |
|
fzyzcjy
|
d5431ff894
|
Tiny add stuck simulation (#15613)
|
2025-12-22 17:00:18 +08:00 |
|
Qiaolin Yu
|
a92de891b8
|
Split dpsk fp4 4 gpu tests and move the mtp part to real stage b (#15553)
|
2025-12-21 15:51:43 -08:00 |
|
Jincong Chen
|
350fbbf4dc
|
fix ds3.2 nsa backend prefill TBO (#14901)
|
2025-12-21 13:16:46 -08:00 |
|
Junrong Lin
|
bed301a5ac
|
[Feature] Enable return routed experts (#12162)
Co-authored-by: yizhang2077 <1109276519@qq.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-12-21 15:16:43 +08:00 |
|
Ke Bao
|
8fe3e37468
|
Support piecewise cuda graph for dsv3 fp4 (#15531)
|
2025-12-21 14:50:32 +08:00 |
|
Alison Shao
|
9a3bdf2c95
|
[CI] Migrate CUDA Graph tests to test/registered/cuda_graph/ (#15436)
|
2025-12-20 18:55:39 -08:00 |
|
michael-amd
|
2ee6c810b8
|
[AMD] Add TP=8 models to nightly test and make TP=2 test stable (#15296)
|
2025-12-19 13:15:19 -08:00 |
|
Yuwei An
|
9d0347b33a
|
EP Support for Piecewise Cuda Graph (#14164)
Signed-off-by: Oasis-Git <ayw.sirius19@gmail.com>
|
2025-12-20 01:59:27 +08:00 |
|
Zehuan Li
|
b2803ff207
|
[DLLM] Add CI for diffusion LLMs (#14723)
|
2025-12-19 08:54:03 +08:00 |
|
Shangming Cai
|
17e81c750a
|
[PP] Add dynamic chunking PP test (#15395)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2025-12-19 01:09:40 +08:00 |
|
Alison Shao
|
58c840db6b
|
Split test_piecewise_cuda_graph.py to optimize CI resource usage (#15290)
|
2025-12-17 21:20:21 -08:00 |
|
Douglas Yang
|
9e7656be80
|
fix: adjust time for test_epd_disaggregation.py (#15354)
|
2025-12-17 18:30:52 -08:00 |
|
Alison Shao
|
4128d4f5cb
|
[CI] Migrate LoRA tests to test/registered/lora/ (#15176)
|
2025-12-17 13:19:42 -08:00 |
|
b8zhong
|
ffa7e03506
|
[Piecewise CUDA Graph] Support INT8 (#14918)
|
2025-12-17 18:20:57 +08:00 |
|
Alison Shao
|
49237e26fb
|
Fix test_pp_single_node.py estimated time from 800s to 500s (#15291)
|
2025-12-16 16:27:09 -08:00 |
|
Shangming Cai
|
6682475124
|
[CI] Improve flaky 4 GPU test success rate (#15234)
|
2025-12-17 00:02:26 +08:00 |
|
Shangming Cai
|
36fcf71fff
|
[Qwen3-next] Add PD disaggregation support for mamba with extra_buffer (#15180)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com>
|
2025-12-16 14:36:00 +08:00 |
|
XDaoHong
|
4733fcff1f
|
[Feature] npu support enable_torch_compile for torchair backend (#13410)
Co-authored-by: ZhengdQin <zhengdqin@gmail.com>
|
2025-12-16 09:23:51 +08:00 |
|
Hanming Lu
|
e61dabf5e4
|
[Qwen3-next] support mamba radix cache for overlap scheduler (#14792)
|
2025-12-14 18:54:16 -08:00 |
|
Tianyu Guo
|
9acb21ae27
|
feat: support EPD disaggregation (#12263)
Co-authored-by: liusy58 <liusy58@linux.alibaba.com>
Co-authored-by: ZhengWG <zwg0606@gmail.com>
Co-authored-by: Nicholas <45984215+liusy58@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
|
2025-12-14 22:30:08 +08:00 |
|
Qiaolin Yu
|
729529190d
|
[ci] Move dpsk-r1-fp4 b200 test to stage b (#15084)
|
2025-12-13 23:14:08 -08:00 |
|
Kangyan-Zhou
|
b243154614
|
Fix CI by reverting incorrect metric check logic (#15004)
|
2025-12-12 10:07:38 -08:00 |
|
Alison Shao
|
e59435c34b
|
Add retry logic for scheduled CI tests (#14771)
|
2025-12-11 16:59:58 -08:00 |
|
Vladimir221
|
27032cecd9
|
[Ascend]Support of piecewise graph compilation for prefill on NPU (#12287)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2025-12-11 21:10:07 +08:00 |
|
Yuan Luo
|
03836d85d2
|
[GLM-4.6V] Support Pipeline Parallelism for GLM-4.6V & GLM-4.1V (#14720)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-12-10 16:40:12 +08:00 |
|
b8zhong
|
56e5c07424
|
fix b200 fa4 ci (#14788)
|
2025-12-10 00:03:43 -08:00 |
|
b8zhong
|
b0a25d0913
|
fix b200 ci (#14786)
|
2025-12-09 23:08:41 -08:00 |
|
Ethan (Yusheng) Su
|
0c63fb9420
|
[Feature] Add LoRA support for embedding layers (#14177)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Beichen-Ma <bm685@cornell.edu>
|
2025-12-09 15:53:33 -08:00 |
|
b8zhong
|
55504df2f7
|
Add FP8 Blockwise GEMM Backend Flag --fp8-gemm-backend (#14379)
|
2025-12-09 12:05:56 -08:00 |
|
Xiaoyu Zhang
|
53d170883a
|
Add fuse_marlin_moe test to ci and add new ep test (#14686)
|
2025-12-09 20:17:38 +08:00 |
|
Alison Shao
|
9a426fc5ef
|
[CI] Move mistral large 3 basic to nightly (#14622)
|
2025-12-09 00:28:44 -08:00 |
|
Alison Shao
|
e6f0ddda44
|
[CI] Migrate Eagle 1-GPU tests to test/registered/ (#14529)
|
2025-12-09 12:56:36 +09:00 |
|