Ke Bao
|
7ace64d1d8
|
Update mamba env setting (#17566)
|
2026-01-23 11:02:32 +08:00 |
|
YC Tseng
|
04a10c9bc2
|
[AMD] CI - migrate perf test and fix stage-b-test-1-gpu-amd (#17340)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
Co-authored-by: michaelzhang-ai <michaelzhang.ai@users.noreply.github.com>
|
2026-01-22 18:45:05 -08:00 |
|
Alison Shao
|
1e8e0cca2c
|
Update test README with CI registry documentation and 5090/H100 guidance (#17368)
|
2026-01-22 16:29:08 -08:00 |
|
chenxu214
|
5d299c25c0
|
[NPU] bugfix with Kimi-k2 and bge-reranker-v2 model (#17478)
Co-authored-by: amote-i <49533125+amote-i@users.noreply.github.com>
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-01-22 22:02:05 +08:00 |
|
zhangheng
|
f33022d039
|
[RadixTree][3/N Refactor]:Support unified insert/evict params (#17401)
|
2026-01-22 17:36:31 +08:00 |
|
YC Tseng
|
17807caf82
|
[AMD] fix amd ci dpskv32 (#17432)
Co-authored-by: michaelzhang-ai <michaelzhang.ai@users.noreply.github.com>
|
2026-01-21 20:34:24 -08:00 |
|
Piotr Mazurek
|
d6e2b88288
|
Add Liquid Foundation Model (LFM2) (#16890)
|
2026-01-22 11:11:20 +08:00 |
|
Alison Shao
|
9be2a3a9a3
|
Remove test_gpt_oss_4gpu.py from __not_in_ci__ (keep in per-commit-4-gpu) (#17534)
|
2026-01-21 15:31:03 -08:00 |
|
Lianmin Zheng
|
b74a57a8d9
|
[Auto Sync] Update detokenizer_manager.py, io_struct.py, mu... (20260120) (#17442)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Wangfan Fu <wangfan@x.ai>
|
2026-01-21 14:48:32 -08:00 |
|
Alison Shao
|
85d9af51da
|
Temporarily disable flaky test_gpt_oss_4gpu.py on B200 (#17528)
|
2026-01-21 14:01:04 -08:00 |
|
Yi Zhang
|
236772c0e1
|
[RadixTree][2/N Refactor]: swa cache init tiny refactor (#17397)
|
2026-01-21 15:48:30 +08:00 |
|
Kangyan-Zhou
|
be5121b452
|
Fix NSA indexer in the nightly test (#17452)
|
2026-01-20 19:58:13 -08:00 |
|
Baizhou Zhang
|
c3f9c30f99
|
[Minor] Change lora_target_modules to "all" in CI tests (#17386)
|
2026-01-21 11:46:36 +08:00 |
|
Alison Shao
|
823a046e8f
|
Add hybrid parallelism test to nightly CI (#17444)
|
2026-01-20 17:43:50 -08:00 |
|
ympcMark
|
f7a5e425c3
|
[3/N] Achieve fault tolerance at the DP level (#11657)
Co-authored-by: UNIDY <unidy2002@outlook.com>
Co-authored-by: Hank Han <hanhan7630@outlook.com>
|
2026-01-20 18:47:08 +08:00 |
|
Michael
|
a3addd6203
|
[AMD] Add DeepSeek-V3.2 and VLMs model in nightly tests (#17179)
Co-authored-by: michaelzhang-ai <michaelzhang-ai@users.noreply.github.com>
Co-authored-by: YC Tseng <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-01-19 20:31:56 -08:00 |
|
Alison Shao
|
7e40d52635
|
Move test_autoround.py to stage-b-test-large-1-gpu suite (#17336)
|
2026-01-19 14:36:19 -08:00 |
|
Hudson Xing
|
c1282da236
|
fix(ci): apply MMMU retry logic to all affected test files (#17329)
|
2026-01-19 10:00:29 -08:00 |
|
shuwenn
|
71279e31f7
|
[CI] fix test_vlm_models.py (#17049)
|
2026-01-19 10:00:00 -08:00 |
|
YC Tseng
|
1a053a810c
|
[AMD] CI - add partitions for stage-b-test-small-1-gpu-amd (#17345)
|
2026-01-19 08:16:07 -08:00 |
|
Bingxu Chen
|
2ea02f0642
|
[AMD CI] Migrate and Add More Testcases (#17116)
Co-authored-by: yctseng0211 <yctseng@amd.com>
|
2026-01-19 08:07:39 -08:00 |
|
zhangheng
|
20b0523eca
|
[RadixTree][1/N Refactor]: Support unified match_prefix params (#17142)
Co-authored-by: yizhang2077 <1109276519@qq.com>
Co-authored-by: pansicheng <sicheng.pan.chn@gmail.com>
|
2026-01-19 22:39:40 +08:00 |
|
b8zhong
|
f374623fa9
|
[Refactor] Set fp4-gemm-backend=auto on SM100 and rename fp4-gemm-backend with flashinfer_ prefix (#17309)
|
2026-01-19 20:09:07 +08:00 |
|
Alison Shao
|
8916b9d080
|
Migrate performance, accuracy, and quantization tests to CI registry (#17177)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-01-18 23:25:24 -08:00 |
|
Kartik Ramesh
|
5836324c55
|
KV Cache Events with Attention DP bug fix (#16030) (#16412)
|
2026-01-19 14:13:48 +08:00 |
|
Gaoji Liu
|
858a4d659b
|
support new qwen3_coder_detector (#16744)
Co-authored-by: liugaoji.lgj <liugaoji.lgj@alibaba-inc.com>
|
2026-01-18 21:16:27 -08:00 |
|
Lee Nau
|
84c8390514
|
Use dsv3 optimized routing fused_topk_deepseek instead of moe_fused_gate (#15347)
|
2026-01-19 11:50:16 +08:00 |
|
Glen Liu
|
ad1b4e4728
|
[Feature] overlap LoRA weight loading with compute (#15512)
|
2026-01-19 10:43:17 +08:00 |
|
b8zhong
|
4df74eb576
|
[Refactor] Add -fp4-gemm-backend to replace SGLANG_FLASHINFER_FP4_GEMM_BACKEND (#16534)
Co-authored-by: Vincent Zhong <207368749+vincentzed@users.noreply.github.com>
|
2026-01-18 23:25:46 +08:00 |
|
Ke Bao
|
f3a7c7dcd9
|
Move radix cache related tests (#17295)
|
2026-01-18 17:47:32 +08:00 |
|
Ke Bao
|
1fe0c82f67
|
Add kl test for swa radix cache (#17281)
|
2026-01-18 16:19:57 +08:00 |
|
Zehuan Li
|
d2c863878c
|
[DLLM] Implement initial dynamic batching for diffusion LLM (#14883)
|
2026-01-17 16:48:15 +08:00 |
|
Yi Zhang
|
737a1183d6
|
[BUGFIX] fix radix cache memory consumption to avoid OOM (#17191)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2026-01-17 16:47:37 +08:00 |
|
StonyPort
|
3355b6e21b
|
feat: add request queued timeout (#17143)
Co-authored-by: qiuxuan.lzw <qiuxuan.lzw@alibaba-inc.com>
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
|
2026-01-16 17:55:09 +08:00 |
|
YC Tseng
|
968c4f55b1
|
[AMD] Enable DeepseekV3.2 test for AMD CI (#16934)
|
2026-01-15 21:58:46 -08:00 |
|
Lianmin Zheng
|
e7dc85c50b
|
Fix grammar sync across TP ranks (#17100)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2026-01-15 18:38:01 -08:00 |
|
Qiaolin Yu
|
e3a95077bc
|
Add dpsk-r1-fp4 in nightly perf ci (#16882)
|
2026-01-15 16:13:35 -08:00 |
|
Alison Shao
|
146b5fcc84
|
[CI] Reorganize stage-b 1-GPU tests for 5090 compatibility (#16826)
|
2026-01-15 15:23:35 -08:00 |
|
Alison Shao
|
69822c7271
|
Disable unit-test-deepep-8-gpu (#17176)
|
2026-01-15 15:12:45 -08:00 |
|
Douglas Yang
|
655d2c7c2a
|
fix: adding matrix partitioning for h200 and b200 nightly tests (#17091)
|
2026-01-15 11:23:19 -08:00 |
|
Shangming Cai
|
4c59782e0f
|
Fix hybrid attention PD Disaggregation test (#17099)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-01-15 23:38:58 +08:00 |
|
Bingxu Chen
|
98096b5e02
|
[AMD CI] migrate and re-enable CI tests to new CI registry (#16949)
Co-authored-by: yctseng0211 <yctseng@amd.com>
|
2026-01-14 21:25:25 -08:00 |
|
Ke Bao
|
7f8a58fffb
|
Refactor prefix cache type checking (#17028)
|
2026-01-15 11:28:13 +08:00 |
|
Артем Савкин
|
424a380077
|
[NPU] NPU quantization refactoring & more quantization formats support (#14504)
Co-authored-by: TamirBaydasov <mr.jeijy@gmail.com>
Co-authored-by: Tamir Baydasov <41994229+TamirBaydasov@users.noreply.github.com>
Co-authored-by: Савкин Артем <savkinartem@MacBook-Air-Viktoria.local>
Co-authored-by: Edward Shogulin <edward.shogulin@gmail.com>
|
2026-01-15 04:25:15 +08:00 |
|
Douglas Yang
|
aa2b4f7661
|
fix: renaming test file and job names + skip blocking llama4 nightly (#16971)
|
2026-01-14 09:57:59 -08:00 |
|
shuwenn
|
de94d793ad
|
feat: support qwen3(-VL) rerank scoring&chat template (#16403)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-01-15 00:45:46 +08:00 |
|
fxmarty-amd
|
5af84c8af5
|
[AMD][Quantization] Add int4fp8_moe online quantization on ROCm (#7392)
Co-authored-by: Dehua Tang <dehtang@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: YC Tseng <yctseng@amd.com>
|
2026-01-14 01:44:40 -08:00 |
|
Netanel Haber
|
e75299a111
|
Fix issues/16714: Revert comment out of tl.debug_barrier() in causal_conv1d_triton (#16899)
Co-authored-by: Yi Zhang <1109276519@qq.com>
|
2026-01-14 17:26:48 +08:00 |
|
Michael
|
b025cff441
|
[AMD] Add AMD CI registration (1-gpu unit test) to nightly CI. (#16941)
|
2026-01-13 23:41:45 -08:00 |
|
Alison Shao
|
c5e363e8e0
|
test: split Qwen3 Next tests and disable PCG tests due to intermittent failures (#16989)
|
2026-01-13 20:18:42 -08:00 |
|