Commit Graph
67 Commits
Author SHA1 Message Date
HAI 3cafb02e5b AMD/ROCm: update daily image repo to lmsysorg/sglang-rocm (#20720) 2026-03-16 20:04:52 -07:00
HAI a0f3361023 Update (#19351) 2026-02-25 11:37:07 -08:00
HAI 6a999dbdf8 [AMD] ENV flags tuning and cleanup (#19176) 2026-02-22 22:40:00 -08:00
HAIandLiangsheng Yin b2573fe426 Upd: CODEOWNERS (#19055)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-02-21 07:51:53 +08:00
HAI 0215d47007 [AMD] ROCm7.2: Add /sgl-workspace/aiter to PYTHONPATH (#18972) 2026-02-18 02:21:39 -08:00
HAI 934b36693c Reasoning models fix docs (#18963) 2026-02-17 23:05:55 -08:00
HAI 8bb1037796 ROCm use rotary_embedding from sgl-kernel (#18920) 2026-02-17 03:00:37 -08:00
HAI b158f5d4a2 Revert "[AMD] Fix RotaryEmbedding crash on AMD/ROCm (regression from #17934)" (#18922) 2026-02-17 01:07:50 -08:00
HAI f4417475b8 Build ROCm7.2 Image with latest AITER v0.1.10.post3 (#18741) 2026-02-12 14:30:13 -08:00
HAI 750ad0d290 [AMD] enable MoRI to release and nightly builds (#18101) 2026-02-02 00:28:30 -08:00
HAI 65d376b491 aiter update to v0.1.6.post1 (#12004) 2025-10-22 23:53:05 -07:00
HAI d500eb9173 aiter v0.1.5.post2 (#10563) 2025-09-17 22:10:45 -07:00
HAI 44426e54be Update REVIEWERS (#9063) 2025-08-11 11:04:39 -07:00
b819381fec AITER backend extension and workload optimizations (#6838)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Hubert Lu <Hubert.Lu@amd.com>
2025-06-05 23:00:18 -07:00
HAI 183d9f969c DeepSeek: enable none block-quant FP8 quantizations (#6638) 2025-05-27 09:06:40 -07:00
HAI 5c0b38f369 aiter attention-backend (default enabled on AMD/ROCm) (#6381) 2025-05-20 22:52:41 -07:00
HAI 6317c5c61f Address performance regression: disable multiple streams on ROCm (#6412) 2025-05-19 21:16:20 -07:00
HAI d364b9b0f2 ROCm: update AITER (#5816) 2025-04-28 11:01:20 -07:00
HAI b0feda090c Revert "Support aiter RMSNorm in AMD" (#5646) 2025-04-22 15:20:24 -07:00
HAI 8879944800 ROCm/AITER CK_MoE: update 2-stage kernels & support both Activations (#5228) 2025-04-10 18:19:57 -07:00
HAI d050df368c ROCm sgl-kernel: compatible to later torch (#5167) 2025-04-10 09:18:36 -07:00
HAIandLianmin Zheng 819924748a Fix refactor error - fp8.py (#5106)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2025-04-07 00:34:08 -07:00
HAIandlinsun12 0beea4503f ROCm: Flex Attention Enablement with custom backends (#4178)
Co-authored-by: linsun12 <linsun12@amd.com>
2025-03-07 04:38:53 -08:00
HAI 13bc39c5d6 ROCm: enable trillion-parameter MoE models with INT4-FP8 single node (#4152) 2025-03-06 15:33:02 -08:00
HAI 71ab0dabe0 Fix the moe padding conditional logic (#4081) 2025-03-05 10:56:51 -08:00
HAI 51d25405a7 ROCm: update aiter and its usage to fused moe (bloat16, fp8, fp8 block-quant) (#4053) 2025-03-04 03:00:46 -08:00
HAI 5c54ef0352 AMD/ROCm: update AITER repo to ROCm/aiter (#3747) 2025-02-21 00:18:08 -08:00
HAI 6252ade985 revert BLOCK and num_warps on HIP (#3722) 2025-02-20 23:30:18 +08:00
HAI d973c78e79 ROCm docker: triton update (#3584) 2025-02-14 10:26:32 -08:00
HAI d81ac4434e MI30x: More graph captures for larger batch sizes and concurrencies (#3420) 2025-02-12 03:04:38 +08:00
HAI 2c1a695ff1 ROCm: sgl-kernel enablement starting with sgl_moe_align_block (#3287) 2025-02-04 21:44:44 +08:00
HAI 566d61d90f ROCm: bump 6.3.0 (#3259) 2025-02-03 04:13:40 +08:00
HAI 17dbf976c5 update ENV to ROCm dockers (#3248) 2025-02-01 17:27:43 +08:00
HAI 6d08ce2aa9 Use Optional with None default (#2770) 2025-01-07 01:35:08 -08:00
HAI c5210dfa38 AMD DeepSeek_V3 FP8 Numerical fix (#2667) 2024-12-30 21:31:12 +08:00
HAI e6f523b5f2 fix typo in python/sglang/srt/layers/quantization/fp8.py (#2655) 2024-12-29 23:45:02 -08:00
HAI 30828e7192 AMD: set weights and scaling numbers properly for block FP8 (#2637) 2024-12-29 03:23:39 -08:00
HAI 7722c11c1d Regression fix to AMD/ROCm from recent change (#2606) 2024-12-26 20:22:14 -08:00
HAI 95f93f493a Fp8 MoE optimizations on AMD (#2388) 2024-12-07 21:18:26 +08:00
HAI b2986d7aa5 Adding SGLang FP8 Utils (#2348) 2024-12-04 03:01:33 -08:00
HAI 0639bf15d1 ROCm Container: set SGLANG_SET_CPU_AFFINITY=1 (#2328) 2024-12-02 23:20:33 -08:00
HAI 69e2d4fb66 Relax to include more AMD GPUs (#2319) 2024-12-02 19:05:58 -08:00
HAI c54bda300a Use rocminfo instead of rocm-smi for more OS/WSL support (#2310) 2024-12-02 00:15:45 -08:00
HAI b79fffdcb5 Update Install Method 2. From source (#2232) 2024-11-27 22:46:55 -08:00
HAI cd51758fad Rename tuned MI300X config files for fused_moe_triton (#2228) 2024-11-27 21:18:51 -08:00
HAI 10189d08dd [Performance]: Process affinity to CPU cores with multiple sockets support (#2171) 2024-11-25 14:57:32 -08:00
HAI f35cb46cc3 ROCm: Fix MoE padding for none FP8 cases (#2111) 2024-11-21 12:23:21 -08:00
HAI e57c3e12b8 Use native fp8 format on MI300X (#2094) 2024-11-19 14:06:29 -08:00
HAI 2ffe0a7363 Add get_amdgpu_memory_capacity() (#2049) 2024-11-15 22:51:48 -08:00
HAI e5c6715003 Fix core (MI300X) with --enable-overlap (#2048) 2024-11-15 21:24:42 -08:00
HAI b275ce0043 Github runner instructions for AMD (#2031) 2024-11-13 23:57:18 -08:00
HAI 087ab83223 [Performance, Triton] Optimize over mask compute to tl.load in fused_moe_kernel (#1980) 2024-11-10 18:54:43 -08:00
HAI f9a377f650 [Release, ROCm] release ROCm docker build for AMD MI GPUs (#1957) 2024-11-08 00:14:15 -08:00
HAI d32fba2a4d [ENV, ROCm] update environment settings (#1939) 2024-11-07 18:24:36 -08:00
HAI 67c424cce3 [Performance, Triton Kernel Args] extend_attention, optimize kern args to _fwd_kernel (#1941) 2024-11-07 18:24:02 -08:00
HAI dca87ec348 [Docs] fix 404 - Contributor Guide (#1942) 2024-11-07 16:50:45 +08:00
HAI 3cd2809277 [Docs, ROCm] update install to cover ROCm with MI GPUs (#1915) 2024-11-04 17:40:57 +08:00
HAI d8e9d61f86 [Build, ROCm] Dockerfile.rocm for Instinct GPUs, with package updates (#1861) 2024-10-31 16:38:16 -07:00
HAI 2d4ce1b792 [Performance, Triton Kernel Args] _decode_grouped_softmax_reducev_fwd… (#1845) 2024-10-30 17:33:36 -07:00
HAI 5f65e2b830 [Performance, Hardware] MoE weights padding to AMD MI300x GPUs (#1836) 2024-10-30 12:17:32 -07:00
HAI 54dd3ea122 [FP8 KV Cache, Mixtral] Avoid KeyError at loading pre-quantized FP8 m… (#1835) 2024-10-29 13:58:03 -07:00
HAI 5010e0d2ca [3rdparty, document] Add 3rdparty/amd, with profiling and tuning instructions to be added (#1822) 2024-10-29 10:51:02 -07:00
HAI e11ab79e68 [Performance, hardware] MoE tuning update to AMD MI300x GPUs (#1619) 2024-10-10 22:48:15 -07:00
HAI 4d086719e5 [Bug] Fix decode stats error on output_len 1 (#1585) 2024-10-06 08:09:09 +00:00
HAI e0b5dbcec1 [FP8 KV Cache] Avoid KeyError at loading pre-quantized FP8 model with kv_scale (#1559) 2024-10-03 01:52:26 -07:00
HAI aa2750beb3 [Bugfix] Enable SGLang on AMD GPUs via PyTorch for ROCm (#1419) (#1453) 2024-09-18 02:01:35 -07:00
HAI 3a6e04185b [Feature, Hardware] Enable SGLang on AMD GPUs via PyTorch for ROCm (#1420) 2024-09-17 07:43:52 +00:00