Commit Graph

96 Commits

Author SHA1 Message Date
Bingxu Chen
98096b5e02 [AMD CI] migrate and re-enable CI tests to new CI registry (#16949)
Co-authored-by: yctseng0211 <yctseng@amd.com>
2026-01-14 21:25:25 -08:00
Hubert Lu
8716589826 [AMD][Diffusion] support timestep embedding kernel for AMD GPUs (#16766) 2026-01-12 22:17:07 -08:00
YC Tseng
b1ee75ae7b [AMD] CI - enable test case for amd ci : triton_attention_kernels , torch_compile_moe (#16559) 2026-01-12 00:06:31 -08:00
Alison Shao
17cb3c8e49 Enable /rerun-stage workflow URL lookup for fork PRs (#16851) 2026-01-11 23:05:37 +08:00
sunxxuns
64a31d4b75 [diffusion] amd: fix SGLANG_DIFFUSION_ATTENTION_BACKEND env var for diffusion attention backend selection (#16325)
Co-authored-by: root <root@mi300x8-008.atl1.do.cpe.ice.amd.com>
2026-01-09 20:16:58 +08:00
YC Tseng
4e999404c3 [AMD] Fix CI - unit-test-backend-1-gpu-amd-mi35x and unit-test-backend-2-gpu-amd, stage-b-test-small-1-gpu-amd (#16675)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-01-08 10:16:27 -08:00
YC Tseng
3d51ae18a1 [AMD] Turn on AMD CI if rocm.Dockerfile changed (#16634) 2026-01-07 22:51:53 -08:00
Bingxu Chen
f9c0426692 [AMD CI] re-enable testcases missed when migrating ci test files (#16535)
Co-authored-by: michael-amd <michael.zhang@amd.com>
Co-authored-by: yctseng0211 <yctseng@amd.com>
2026-01-07 22:43:48 -08:00
Hubert Lu
b86bbf841e [AMD] Add 8-GPU MX35X test running DSR1-MXFP4 model for AMD CI (#13602) 2026-01-07 01:43:11 -08:00
YC Tseng
8bce085321 [AMD] CI - add 2 pp test cases to performance-test-2-gpu-amd (#16514) 2026-01-06 23:52:46 -08:00
Alison Shao
90eac38a12 Migrate FP8/TorchAO tests to test/registered/quant/ (#16453) 2026-01-06 18:27:43 -08:00
Alison Shao
3271e0e76d Remove dllm-test-1-gpu-amd job (followup to DLLM migration) (#16589) 2026-01-06 14:27:37 -08:00
Alison Shao
861a35fb6e ci: migrate RL tests to test/registered/rl/ (#16417) 2026-01-05 22:13:13 -08:00
yctseng0211
c371df2f25 [AMD] Fix CI and add retry logic for git clone timeout (#15663)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-01-05 12:37:52 -08:00
Alison Shao
f8411ded6e ci: migrate 1-GPU model tests to test/registered/models/ (#16414) 2026-01-04 18:08:01 -08:00
sunxxuns
8b869e326c [AMD] feat: add DLLM support for AMD GPUs with LLaDA2 testing (#15560) 2026-01-03 10:41:11 +08:00
Alison Shao
17041f4673 Enable /rerun-stage slash command to work on fork PRs (#16128) 2026-01-02 10:36:12 -08:00
Kangyan-Zhou
d65ae0ec7a Use a different concurrency group for release branch testing (#16202) 2025-12-31 09:46:43 -08:00
Baizhou Zhang
98225be6e5 [CI] Fixing release with cut branch workflow (#16153)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
2025-12-30 17:18:02 +08:00
sunxxuns
c69c1c4f07 [CI] Fix AMD CI to exclude multimodal_gen from main_package filter (#15558) 2025-12-21 16:58:17 +08:00
Hubert Lu
51e2eaa458 [AMD] Support fast_topk kernels in sgl-kernel (#15172) 2025-12-19 22:19:09 -08:00
yctseng0211
af780c59f3 [AMD] add unit-test-backend-8-gpu-amd back (#15253)
Co-authored-by: Sai Enduri <saimanas.enduri@amd.com>
2025-12-19 00:59:28 -08:00
yctseng0211
a21aa87ec5 [AMD] Fix and add accuracy-test-2-gpu-amd back (#15415)
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
2025-12-19 00:41:14 -08:00
Yuzhen Zhou
4bf06635fc [diffusion] multi-platform: support diffusion on amd and fix encoder loading on MI325 (#13760)
Co-authored-by: Sabre Shao <sabre.shao@amd.com>
Co-authored-by: Yusheng (Ethan) Su <yushengsu.thu@gmail.com>
Co-authored-by: Hubert Lu <Hubert.Lu@amd.com>
Co-authored-by: xsun <sunxiao04@gmail.com>
2025-12-19 15:38:46 +08:00
Sai Enduri
538e733e08 [AMD CI] Fix typo. (#15229) 2025-12-15 21:47:01 -08:00
Sai Enduri
02d24244e4 [AMD CI] Temporarily disable 2 gpu accuracy test. (#15204) 2025-12-15 20:57:26 -08:00
Alison Shao
4c5074eb78 Add AMD stage support to /rerun-stage command and fix related bugs (#14463) 2025-12-04 21:13:52 -08:00
sunxxuns
5bbd83a2c8 ci: Migrate AMD workflows to new MI325 runners; temporarily disabled failed CI's to be added back (#14226) 2025-12-03 11:33:27 -08:00
Liangsheng Yin
f7be98e113 [CI] fix amd 1 gpu basic test (#13551) 2025-11-19 11:47:41 +08:00
Sai Enduri
9a1a9a4209 [AMD CI] Local cache fallback. (#13452)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-11-18 19:23:45 -08:00
Lianmin Zheng
2e1dbdb258 Update docs (#13519)
Co-authored-by: sglang-bot <sglangbot@gmail.com>
2025-11-18 06:24:58 -08:00
Liangsheng Yin
67071f55a8 [CI] fix triggered by a non-run-ci label (#13393) 2025-11-18 14:24:58 +08:00
Liangsheng Yin
ab63f3c50b [1/N] CI refactor: introduce CI register. (#13345) 2025-11-17 12:21:20 +08:00
alisonshao
9bd511a582 Add model validation for all GPU runners to prevent cache corruption (#13171)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
2025-11-13 10:16:41 -08:00
Mick
307e7a6128 diffusion: fix detected file changes rule in CI (#12943) 2025-11-10 13:37:16 +08:00
Hubert Lu
3694266051 Expand and update test coverage for AMD CI (#10044) 2025-11-04 22:15:13 -08:00
Sai Enduri
1d08653972 [AMD CI] Add image and weights caching. (#11593) 2025-10-14 02:51:35 -07:00
sunxxuns
a57f0e3d56 reverse the amd ci test back to 1200s and split the 8-gpu deepseek job into two. (#11238)
Co-authored-by: root <root@smci350-zts-gtu-e17-15.zts-gtu.dcgpu>
2025-10-06 19:27:57 -04:00
sunxxuns
5e142484e2 [Fix AMD CI] VRAM cleanup (#11174)
Co-authored-by: root <root@smci350-zts-gtu-e17-15.zts-gtu.dcgpu>
2025-10-05 19:03:53 -04:00
Sai Enduri
195a59fe23 Refactor AMD CI. (#11128) 2025-10-01 01:12:28 -07:00
Lianmin Zheng
50dc0c1e9c Run tests based on labels (#10456) 2025-09-15 00:29:20 -07:00
Hubert Lu
91b3555d2d Add tests to AMD CI for MI35x (#9662)
Co-authored-by: Sai Enduri <saimanas.enduri@amd.com>
2025-09-10 12:50:05 -07:00
Lianmin Zheng
05e4787243 [CI] Fix the trigger condition for PR test workflows (#9761) 2025-08-30 15:47:10 -07:00
Hubert Lu
711390a971 [AMD] Support Hierarchical Caching on AMD GPUs (#8236) 2025-08-28 15:27:07 -07:00
Hubert Lu
c6c379ab31 [AMD] Reorganize hip-related header files in sgl-kernel (#9320) 2025-08-18 16:53:44 -07:00
Sai Enduri
740f063035 Fix Custom All Reduce CI job. (#9258) 2025-08-16 16:29:43 -07:00
kk
983aa4967b Fix nan value generated after custom all reduce (#8663)
Co-authored-by: wunhuang <wunhuang@amd.com>
2025-08-15 12:33:54 -07:00
Hubert Lu
9c3e95d98b [AMD] Expand test coverage for AMD CI and enable apply_token_bitmask_inplace_cuda in sgl-kernel (#8268) 2025-08-15 12:32:51 -07:00
Lianmin Zheng
2c7f01bc89 Reorganize CI and test files (#9027) 2025-08-10 12:30:06 -07:00
Lianmin Zheng
67a7d1f699 Create cancel-all-pr-test-runs (#8986) 2025-08-08 15:53:51 -07:00