Hubert Lu
|
93423ff780
|
[AMD] Deprecate ROCm 6.3 artifacts and standardize gfx942 on ROCm 7 (#17785)
|
2026-01-27 15:58:49 -08:00 |
|
 Hubert Luandwufann
|
df42f4d386
|
[AMD] Update dsv3.2 AMD GPU docs and unify ROCm TileLang build (#17783)
Co-authored-by: wufann <715544327@qq.com>
|
2026-01-26 21:10:32 -08:00 |
|
 Hubert Luandwufann
|
afe285f7bd
|
[AMD] enable CUDA graph for NSA backend and fix NSA FP8 fused RMSNorm group quant (#16841)
Co-authored-by: wufann <715544327@qq.com>
|
2026-01-13 17:36:01 -08:00 |
|
Hubert Lu
|
8716589826
|
[AMD][Diffusion] support timestep embedding kernel for AMD GPUs (#16766)
|
2026-01-12 22:17:07 -08:00 |
|
 Hubert LuandHAI
|
d6d5c3fdea
|
[AMD] Clean up vllm dependencies in moe_runner/triton.py (#11349)
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-01-09 00:24:04 -08:00 |
|
Hubert Lu
|
4935344fcd
|
[AMD] Fix aiter page-size handling, DeepSeek MLA tuple inputs, and HiCache/FA3 decode-backend override (#16531)
|
2026-01-07 21:14:32 -08:00 |
|
Hubert Lu
|
b86bbf841e
|
[AMD] Add 8-GPU MX35X test running DSR1-MXFP4 model for AMD CI (#13602)
|
2026-01-07 01:43:11 -08:00 |
|
Hubert Lu
|
51e2eaa458
|
[AMD] Support fast_topk kernels in sgl-kernel (#15172)
|
2025-12-19 22:19:09 -08:00 |
|
Hubert Lu
|
2ce2377740
|
[AMD] Fix AITER_MXFP4_MOE_SF setting for gfx950 (#13239)
|
2025-11-13 17:03:41 -08:00 |
|
 
|
e4b2937017
|
[AMD] Add AITER Custom All-Reduce (#13102)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2025-11-12 21:53:44 -08:00 |
|
Hubert Lu
|
36d147121f
|
[AMD] Apply AITER_MXFP4_MOE_SF=1 only to gfx950 in aiter build (#13092)
|
2025-11-11 12:26:09 -08:00 |
|
Hubert Lu
|
3694266051
|
Expand and update test coverage for AMD CI (#10044)
|
2025-11-04 22:15:13 -08:00 |
|
Hubert Lu
|
01a26544a3
|
[AMD] Add Tilelang and Fast Hadamard Transform builds to Dockerfile.rocm (#11114)
|
2025-09-30 20:00:37 -07:00 |
|
Hubert Lu
|
fe68c1486f
|
Fix errors of hicache kernels in sgl-kernel for ROCm (#10339)
|
2025-09-11 14:54:34 -07:00 |
|
 Hubert LuandSai Enduri
|
91b3555d2d
|
Add tests to AMD CI for MI35x (#9662)
Co-authored-by: Sai Enduri <saimanas.enduri@amd.com>
|
2025-09-10 12:50:05 -07:00 |
|
Hubert Lu
|
2c562fd2d0
|
Fix Llama 4 with MXFP4 dynamic quant on MI35x (#9993)
|
2025-09-04 00:48:58 -07:00 |
|
Hubert Lu
|
711390a971
|
[AMD] Support Hierarchical Caching on AMD GPUs (#8236)
|
2025-08-28 15:27:07 -07:00 |
|
Hubert Lu
|
f445a1d9a3
|
[AMD] Fix Llama 4 FP8 accuracy issues on MI300X (#7699)
|
2025-08-22 13:13:45 -07:00 |
|
Hubert Lu
|
704ced1b2e
|
[AMD] Remove the deprecated C10_WARP_SIZE (#9356)
|
2025-08-21 18:16:35 -07:00 |
|
Hubert Lu
|
c6c379ab31
|
[AMD] Reorganize hip-related header files in sgl-kernel (#9320)
|
2025-08-18 16:53:44 -07:00 |
|
Hubert Lu
|
9c3e95d98b
|
[AMD] Expand test coverage for AMD CI and enable apply_token_bitmask_inplace_cuda in sgl-kernel (#8268)
|
2025-08-15 12:32:51 -07:00 |
|
 
|
af4b9bae95
|
[AMD] Add silu_and_mul, gelu_and_mul, gelu_tanh_and_mul, and gelu_quick kernels for AMD GPUs (#7135)
Co-authored-by: yiakwy-xpu-ml-framework-team <961186938@qq.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2025-07-24 23:44:28 -07:00 |
|
Hubert Lu
|
e50109f2ed
|
[AMD] Remove vllm's scaled_fp8_quant and moe_sum when SGLANG_USE_AITER=1 (#7484)
|
2025-07-21 17:33:19 -07:00 |
|
Hubert Lu
|
7750b91ca8
|
[AMD] Add triton awq_dequantize kernel to support AWQ on ROCm (#7661)
|
2025-07-18 14:27:25 -07:00 |
|
Hubert Lu
|
e00715eb66
|
[AMD] Add test_fused_moe.py and test_rope_rocm.py to AMD CI (#5246)
|
2025-07-06 01:47:16 -07:00 |
|
Hubert Lu
|
b116b21a46
|
[AMD] Temporarily disable test_no_overlap_scheduler and test_vision_chunked_prefill (#7717)
|
2025-07-02 12:39:18 -07:00 |
|
Hubert Lu
|
3b3f1e3aeb
|
[AMD] Add unit-test-sgl-kernel-amd to AMD CI (#7539)
|
2025-06-29 15:50:09 -07:00 |
|
Hubert Lu
|
4740288303
|
[AMD] Add more tests to per-commit-amd (#6926)
|
2025-06-08 01:08:37 -07:00 |
|
Hubert Lu
|
198b9056d1
|
[AMD] Fix Llama 4 Scout and Maverick accuracy issues on MI300X (#6274)
|
2025-05-14 22:07:29 +00:00 |
|
Hubert Lu
|
2a936a841e
|
[AMD] switch to custom allreduce regardless of MSCCL setting on ROCm (#6097)
|
2025-05-08 13:46:58 -07:00 |
|
Hubert Lu
|
afb752bcbe
|
[AMD] Fix missing per_token_group_quant_fp8 for ROCm (#5140)
|
2025-04-07 22:38:25 -07:00 |
|
Hubert Lu
|
9cf4077294
|
Enable custom AR for AMD GPUs and maintain it in sgl-kernel (#3406)
|
2025-03-02 15:19:06 -08:00 |
|
Hubert Lu
|
f8b28e461a
|
Add CPU affinity setting to latency benchmark (#3085)
|
2025-01-25 23:52:05 -08:00 |
|