Alison Shao
|
f4ec6f8e17
|
ci: migrate remaining spec/eagle tests to test/registered/spec/ (#15800)
|
2025-12-29 14:00:00 -08:00 |
|
Cheng Wan
|
6f9d0a89a0
|
[scheduler] fix: correcting extend_logprob_start_len calculation (#15922)
|
2025-12-28 14:57:04 -08:00 |
|
lif
|
5969be2f06
|
Apply fixture-kit mode to MMMUVLMMixin (#15615)
|
2025-12-28 17:22:05 +08:00 |
|
Cheng Wan
|
c457aad54a
|
Update test parameters for deepep_large test (#16001)
|
2025-12-28 00:58:19 -08:00 |
|
Liangsheng Yin
|
c7e7bfa32c
|
Add EAGLE3 test with MMLU dataset. (#15945)
|
2025-12-28 16:07:58 +08:00 |
|
Liangsheng Yin
|
bf90ea9c5b
|
Unify spec v2's naming manner. (#15990)
|
2025-12-28 14:14:52 +08:00 |
|
Lianmin Zheng
|
183b65190a
|
Clean up logging (#15919)
|
2025-12-27 15:27:12 -08:00 |
|
Ke Bao
|
faecd37ed4
|
Add Mimo-v2-flash model to ci test (#15887)
|
2025-12-27 14:18:08 +08:00 |
|
Liangsheng Yin
|
9ad546d7e8
|
Tiny cleanup the models' name in test_utils (#15920)
|
2025-12-27 14:13:23 +08:00 |
|
Alison Shao
|
72a980c6d5
|
ci: migrate MLA tests to test/registered/mla/ (#15798)
|
2025-12-24 23:29:30 -08:00 |
|
satyamk7054
|
38dd4fbb66
|
Add overlap scheduling for embeddings code path (#14032)
Co-authored-by: Satyam Kumar <satyamk@linkedin.com>
|
2025-12-24 18:24:18 -08:00 |
|
fzyzcjy
|
8196998a03
|
Tiny add num retracted tokens metric (#15653)
|
2025-12-24 22:15:23 +08:00 |
|
fzyzcjy
|
fd4a558e71
|
Add metrics for having prefill and decode in different ranks (#15752)
|
2025-12-24 21:35:35 +08:00 |
|
michael-amd
|
e7b09efc0a
|
[AMD] Add AMD Nightly Performance & VLMs Accuracy Tests (#15500)
|
2025-12-23 19:03:27 -08:00 |
|
Liangsheng Yin
|
0d0367e9d0
|
Tiny fix CI (#15696)
|
2025-12-24 01:51:08 +08:00 |
|
Liangsheng Yin
|
bd572360f3
|
Tiny apply gsm8k mixin to ngram test (#15606)
|
2025-12-24 01:30:26 +08:00 |
|
fzyzcjy
|
6a5764a719
|
Super tiny add test_soft_watchdog to nightly (#15692)
|
2025-12-23 22:46:15 +08:00 |
|
Alison Shao
|
ac42797cf7
|
[CI] Enable retry logic for flaky CI tests (#14983)
|
2025-12-22 22:23:42 -08:00 |
|
Alison Shao
|
883747ced1
|
[CI] Migrate Attention Backend tests to test/registered/attention/ (#15563)
|
2025-12-22 22:17:52 -08:00 |
|
SeanWeiSean
|
7759716786
|
bugfix[schedule]: Refactor sort method and add related UT (#13576)
Co-authored-by: Yuxuan Wei <w1300012920@pku.edu.cn>
Co-authored-by: Yuxuan Wei <w1300012920@gmail.com>
Co-authored-by: Yuxuan Wei🚚 <yuxwei@microsoft.com>
|
2025-12-23 03:13:57 +08:00 |
|
fzyzcjy
|
d5431ff894
|
Tiny add stuck simulation (#15613)
|
2025-12-22 17:00:18 +08:00 |
|
Liangsheng Yin
|
beae3f961c
|
Adapt fixture-kit to gsm8k mixin (#15599)
|
2025-12-22 14:19:35 +08:00 |
|
Baizhou Zhang
|
468931b572
|
[Tiny]Move deepseek fp4 cutlass moe test to per-commit test (#15565)
|
2025-12-21 18:08:07 -08:00 |
|
Qiaolin Yu
|
a92de891b8
|
Split dpsk fp4 4 gpu tests and move the mtp part to real stage b (#15553)
|
2025-12-21 15:51:43 -08:00 |
|
Jincong Chen
|
350fbbf4dc
|
fix ds3.2 nsa backend prefill TBO (#14901)
|
2025-12-21 13:16:46 -08:00 |
|
Junrong Lin
|
bed301a5ac
|
[Feature] Enable return routed experts (#12162)
Co-authored-by: yizhang2077 <1109276519@qq.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-12-21 15:16:43 +08:00 |
|
Ke Bao
|
8fe3e37468
|
Support piecewise cuda graph for dsv3 fp4 (#15531)
|
2025-12-21 14:50:32 +08:00 |
|
Alison Shao
|
9a3bdf2c95
|
[CI] Migrate CUDA Graph tests to test/registered/cuda_graph/ (#15436)
|
2025-12-20 18:55:39 -08:00 |
|
mlmz
|
1f1f05a85e
|
vlm: refactor engine vlm params and support processor output as input (#14091)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: zhaochenyang20 <zhaochenyang20@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: BenYao21 <cyao22@asu.edu>
Co-authored-by: minleminzui <minleminzui@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
|
2025-12-20 18:31:24 +08:00 |
|
michael-amd
|
2ee6c810b8
|
[AMD] Add TP=8 models to nightly test and make TP=2 test stable (#15296)
|
2025-12-19 13:15:19 -08:00 |
|
Yuwei An
|
9d0347b33a
|
EP Support for Piecewise Cuda Graph (#14164)
Signed-off-by: Oasis-Git <ayw.sirius19@gmail.com>
|
2025-12-20 01:59:27 +08:00 |
|
yctseng0211
|
af780c59f3
|
[AMD] add unit-test-backend-8-gpu-amd back (#15253)
Co-authored-by: Sai Enduri <saimanas.enduri@amd.com>
|
2025-12-19 00:59:28 -08:00 |
|
Zehuan Li
|
f6c9db4bc4
|
[DLLM] Fix dLLM regression (#15371)
|
2025-12-19 10:55:46 +08:00 |
|
Zehuan Li
|
b2803ff207
|
[DLLM] Add CI for diffusion LLMs (#14723)
|
2025-12-19 08:54:03 +08:00 |
|
sunxxuns
|
e0963a6cb1
|
[AMD] Clear pre-built AITER kernels and warmup to prevent segfaults and test timeouts (#15318)
|
2025-12-18 14:00:09 -08:00 |
|
Shangming Cai
|
17e81c750a
|
[PP] Add dynamic chunking PP test (#15395)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2025-12-19 01:09:40 +08:00 |
|
Lianmin Zheng
|
d1f0063262
|
Clean up __init__ function of the scheduler and event loop for PD (#15298)
|
2025-12-18 01:35:14 -08:00 |
|
Alison Shao
|
58c840db6b
|
Split test_piecewise_cuda_graph.py to optimize CI resource usage (#15290)
|
2025-12-17 21:20:21 -08:00 |
|
Douglas Yang
|
9e7656be80
|
fix: adjust time for test_epd_disaggregation.py (#15354)
|
2025-12-17 18:30:52 -08:00 |
|
Alison Shao
|
4128d4f5cb
|
[CI] Migrate LoRA tests to test/registered/lora/ (#15176)
|
2025-12-17 13:19:42 -08:00 |
|
Kangyan-Zhou
|
011d8d8970
|
Reserve more memory for DeepSeekOCR model and adjust server start timeout for DeepGEMM to reduce flakiness (#15277)
|
2025-12-17 13:13:05 -08:00 |
|
Liangsheng Yin
|
0c00220795
|
tiny unify environ usage (#15335)
|
2025-12-17 23:31:43 +08:00 |
|
b8zhong
|
ffa7e03506
|
[Piecewise CUDA Graph] Support INT8 (#14918)
|
2025-12-17 18:20:57 +08:00 |
|
Xuchun Shang
|
45a959d3e9
|
[PP] Add pp support for Qwen3-VL (#12333)
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Signed-off-by: Kun(llfl) <i@imux.top>
Signed-off-by: Kun(llfl) <llfl@linux.alibaba.com>
Co-authored-by: kun-llfl <i@imux.top>
Co-authored-by: Kun(llfl) <llfl@linux.alibaba.com>
|
2025-12-17 16:03:58 +08:00 |
|
elvischenv
|
435d1c83c1
|
[Perf] Enable Flashinfer autotune by default (#14357)
|
2025-12-16 23:01:39 -08:00 |
|
Douglas Yang
|
46ad4b986d
|
fix: moving decorator to header (#15297)
|
2025-12-16 17:05:45 -08:00 |
|
Douglas Yang
|
da58df6b3d
|
fix: skipping TestEPDDisaggregationOneEncoder test (#15292)
|
2025-12-16 16:31:20 -08:00 |
|
Alison Shao
|
49237e26fb
|
Fix test_pp_single_node.py estimated time from 800s to 500s (#15291)
|
2025-12-16 16:27:09 -08:00 |
|
amysaq2023
|
ccc8f3b266
|
support non disturbing remote instance weight loader v2 (#14997)
Signed-off-by: Anqi Shen <amy.saq@antgroup.com>
|
2025-12-16 14:39:56 -08:00 |
|
Shangming Cai
|
6682475124
|
[CI] Improve flaky 4 GPU test success rate (#15234)
|
2025-12-17 00:02:26 +08:00 |
|