Commit Graph

  • d973c78e79 ROCm docker: triton update (#3584) HAI 2025-02-14 10:26:32 -08:00
  • 6ce6eabbcc Copy config files for MI300X to support in virtualized environments (#3505) Jesse Lopez 2025-02-14 09:23:32 -08:00
  • 4e23c961e8 docs: update install (#3581) Yineng Zhang 2025-02-14 18:54:50 +08:00
  • 3efbdf68b9 fix sgl-kernel codestyle (#3563) Xiaoyu Zhang 2025-02-14 18:05:52 +08:00
  • 6cc309557a Add support for OpenAI API o1 model (#3363) Chuyue Sun 2025-02-13 19:43:00 -08:00
  • 31eec35ba8 fix doc (#3558) Yineng Zhang 2025-02-14 10:11:31 +08:00
  • ac963be234 update flashinfer-python (#3557) Yineng Zhang 2025-02-14 09:52:56 +08:00
  • e0b9a423c8 chore: bump v0.4.3 (#3556) Yineng Zhang 2025-02-14 09:43:14 +08:00
  • e082142519 chore: bump 0.0.3.post6 sgl-kernel (#3555) Yineng Zhang 2025-02-14 08:55:15 +08:00
  • 70f894b810 feat: support flashinfer mla attention for deepseek v3 (#3550) Yineng Zhang 2025-02-14 08:50:14 +08:00
  • 368de3661e Update install docs (#3553) simveit 2025-02-13 22:42:51 +01:00
  • 20de05a753 update README (#3543) Yineng Zhang 2025-02-13 17:22:11 +08:00
  • f076328bb7 fix moe_align_kernel shm init not sync bug (#3534) Xiaoyu Zhang 2025-02-13 16:47:00 +08:00
  • bf2a70872e Update DeepSeek V3 Doc (#3541) Jhin 2025-02-13 01:15:37 -06:00
  • 871a4aa1bf [ROCm] Add ROCm tuning configs for AMD Instinct MI325X. (#3536) Wen-Heng (Jack) Chung 2025-02-12 22:09:36 -06:00
  • 98eecbda54 integrate blockwise fp8 kernel (#3529) yizhang2077 2025-02-13 04:39:33 +08:00
  • 4430c0a513 chore: bump 0.0.3.post5 sgl-kernel (#3530) Yineng Zhang 2025-02-13 01:51:46 +08:00
  • 640363ad20 support blockwise fp8 matmul kernel (#3267) yizhang2077 2025-02-13 01:49:33 +08:00
  • 8616357a97 Fix deepseek awq v3 (#3450) Liangsheng Yin 2025-02-12 22:09:52 +08:00
  • 8adbc78b30 added llama and cleaned up (#3503) Zachary Streeter 2025-02-12 04:48:30 -06:00
  • 45e3a7bc41 use sgl_per_token_group_quant_fp8 kernel (#3493) Xiaoyu Zhang 2025-02-12 18:40:42 +08:00
  • b96e92e6e6 chore: bump 0.0.3.post4 sgl-kernel (#3523) Yineng Zhang 2025-02-12 17:28:36 +08:00
  • 693c2600e0 refine deepseek_v3 launch server doc (#3522) Xiaoyu Zhang 2025-02-12 17:27:07 +08:00
  • ced680663c doc: Support a new vLM (#3405) Mick 2025-02-12 16:43:14 +08:00
  • b8318aec48 Make NCCL NVLS configurable (#3502) Ata Fatahi 2025-02-11 14:25:06 -05:00
  • 2f48221033 docs: update install Yineng Zhang 2025-02-12 03:13:31 +08:00
  • d81ac4434e MI30x: More graph captures for larger batch sizes and concurrencies (#3420) HAI 2025-02-11 11:04:38 -08:00
  • 2491cc928d add deepseek-v3 amd docker command (#3495) Zachary Streeter 2025-02-11 13:03:08 -06:00
  • 67c5de9286 fix router typo (#3496) Didier Durand 2025-02-11 20:00:57 +01:00
  • 1e2cf2b541 fix server_arguments typo (#3499) Didier Durand 2025-02-11 19:59:53 +01:00
  • 9490d15772 fix supported_models Qwen typo (#3498) Didier Durand 2025-02-11 19:59:18 +01:00
  • eefcbdd353 fix deepseek_v3 typo (#3497) Didier Durand 2025-02-11 19:58:36 +01:00
  • 7e6d5fc694 Support Eagle cuda graph for Triton backend (#3500) Ke Bao 2025-02-12 02:27:45 +08:00
  • cadd5dbe6a Tune MI300X fused MoE Triton kernel JSON config. (#3492) Wen-Heng (Jack) Chung 2025-02-11 12:27:25 -06:00
  • bb418ced80 optimize per token group quant fp8 (#3490) Xiaoyu Zhang 2025-02-11 22:19:05 +08:00
  • fdf04a1426 [ROCm] Add ROCm tuning config to block gemm and Re-tune for AMD Radeon Graphics (#3418) yigex 2025-02-11 15:55:04 +08:00
  • 5f0e7de339 [Feat] Return hidden states (experimental) (#3364) Jackmin801 2025-02-10 15:54:37 -08:00
  • 2f47d710ae refine some typo (#3473) Xiaoyu Zhang 2025-02-10 23:35:44 +08:00
  • 4fe92bfca5 fix mla test (#3469) Yineng Zhang 2025-02-10 21:12:00 +08:00
  • d23cb9a01e [Eagle] reduce one draft forward (#3468) Ying Sheng 2025-02-10 04:21:49 -08:00
  • 2d61132374 Support Eagle2 for Triton backend (#3466) Ke Bao 2025-02-10 20:00:42 +08:00
  • cddb1cdf8f chore: bump v0.4.2.post4 (#3459) Yineng Zhang 2025-02-10 14:12:16 +08:00
  • fa1b40e00d use nvcr.io/nvidia/tritonserver:24.04-py3-min as base image (#3457) Yineng Zhang 2025-02-10 13:52:33 +08:00
  • c45cab1c00 [Fix] Fix accuracy bug and refactor codes for lora (#3413) Baizhou Zhang 2025-02-09 21:29:00 -08:00
  • 27c4c9cf52 remove _grouped_size_compiled_for_decode_kernels (#3453) Yineng Zhang 2025-02-10 13:01:21 +08:00
  • 52a492a16e Update contribution_guide.md (#3452) Ying Sheng 2025-02-09 20:53:47 -08:00
  • 36f6fc5093 feat: enable ragged fa3 by default on hopper 12.4+ (#3442) Yineng Zhang 2025-02-10 07:43:01 +08:00
  • d87272750b fix ci (#3441) Yineng Zhang 2025-02-10 04:22:28 +08:00
  • 6239d0b2e7 chore: bump sgl-kernel v0.0.3.post3 (#3440) Yineng Zhang 2025-02-10 04:00:52 +08:00
  • 4cfd3add6d support version in sgl-kernel (#3439) Yineng Zhang 2025-02-10 03:49:52 +08:00
  • 20cf910d8f [docs] Update quantization documentation (#3437) Shi Shuai 2025-02-09 18:39:49 +00:00
  • 0af1d239cb [Docs] Add quantization docs (#3410) Wenxuan Tan 2025-02-09 12:16:21 -06:00
  • 85986bb978 compatible with new outlines (#3435) Yineng Zhang 2025-02-10 01:51:30 +08:00
  • 64c8713573 remove activation dependency in fused_moe (#3433) Yineng Zhang 2025-02-10 01:18:57 +08:00
  • 1646149a83 fix draft cuda graph capture failure (#3431) Yineng Zhang 2025-02-09 23:16:20 +08:00
  • bc72e5bd32 add cuda graph capture failure possible solution (#3430) Yineng Zhang 2025-02-09 22:57:11 +08:00
  • 014cab4dd2 update forward_return_lse (#3425) Yineng Zhang 2025-02-09 20:18:44 +08:00
  • 4d2dbeaca7 remove cutex dependency (#3422) Yineng Zhang 2025-02-09 18:33:20 +08:00
  • 29daf498cd fix cu118 link issue (#3421) Yineng Zhang 2025-02-09 18:16:44 +08:00
  • 6702592d0e [docs] Add multi-node inference example for SLURM in documentation (#3408) Shi Shuai 2025-02-09 05:45:14 +00:00
  • 60abdb3e7c minor: cleanup test_eagle_infer (#3415) Yineng Zhang 2025-02-09 09:34:30 +08:00
  • 7b4e61fff3 [Fix] Fix eagle with disable cuda graph (#3411) Ying Sheng 2025-02-08 16:40:00 -08:00
  • 6222e1c228 add disable cuda graph unit test for eagle 2 (#3412) Yineng Zhang 2025-02-09 08:02:56 +08:00
  • fad315cb8e fix EAGLE 2 non greedy case (#3407) Yineng Zhang 2025-02-09 07:28:34 +08:00
  • f90db8bc07 fix typo Yineng Zhang 2025-02-08 22:16:42 +08:00
  • d8ad597048 Add deepseek-v3 a100 serving example (#3404) Ke Bao 2025-02-08 22:13:52 +08:00
  • 849f58d617 Update fused_moe's benchmark (#3346) GaoYuYang 2025-02-08 21:58:21 +08:00
  • 64480df495 [BUG] fix moe benchmark when bs*seq is small (#3382) yiakwy-xpu-ml-framework-team 2025-02-08 15:39:44 +08:00
  • 4530136e61 Add H20 fp8 w8a8 gemm config (#3386) lukec 2025-02-08 15:36:31 +08:00
  • 0a6f18f068 added amd_configure.md to references (#3275) Zachary Streeter 2025-02-07 10:50:49 -06:00
  • c1f5f99f60 chore: bump v0.4.2.post3 (#3369) Yineng Zhang 2025-02-08 00:20:03 +08:00
  • fa82dfccdd fix EagleVerifyInput (#3378) Yineng Zhang 2025-02-07 22:30:43 +08:00
  • 5da3d21c8b update pr-test ci (#3376) Yineng Zhang 2025-02-07 21:08:35 +08:00
  • f287037673 update sgl-kernel version (#3374) Yineng Zhang 2025-02-07 20:51:06 +08:00
  • f9905d59a8 support speculative decoding kernel in sgl-kernel (#3373) Yineng Zhang 2025-02-07 20:29:51 +08:00
  • 45c87e083f fix undefined symbol cudaGetDriverEntryPointByVersion (#3372) Yineng Zhang 2025-02-07 19:32:45 +08:00
  • 2b1808cec4 update unit test in AMD CI (#3366) Yineng Zhang 2025-02-07 17:25:16 +08:00
  • e868d0b60e update waves_per_eu to 1 (#3356) lizamd 2025-02-06 21:08:06 -08:00
  • 591e751e07 Fix: Runtime error for function calling (#3300) Shi Shuai 2025-02-07 04:52:01 +00:00
  • 40022d075a Feature: Fix the binding error in Llama (#3355) Chayenne 2025-02-06 20:19:24 -08:00
  • 823148e7f0 Docs: Add deepseek usage and add multi-node, credit to lycanlancelot (#3314) Liangjun Song 2025-02-07 14:45:00 +11:00
  • 76ca91dff2 Docs/CI: Enable Fake Finish for Docs Only PR (#3350) Chayenne 2025-02-06 19:33:31 -08:00
  • cdae77b03d optimize moe_align_kernel cuda (#3347) Xiaoyu Zhang 2025-02-07 00:53:46 +08:00
  • adeee15204 fix sgl-kernel build failure on AMD (#3352) Yineng Zhang 2025-02-07 00:35:59 +08:00
  • 6792411e7f [Doc] Add optimization option guide for deepseek v3 (#3349) Ke Bao 2025-02-06 23:28:09 +08:00
  • 7348d9627e add AMD guide for DeepSeek-R1 (#3338) Yineng Zhang 2025-02-06 16:54:40 +08:00
  • 25ed22b685 update pull request template (#3337) Yineng Zhang 2025-02-06 16:48:02 +08:00
  • 200d3b1608 Add sgl-kernel to MI300 CI paths tested. (#3335) saienduri 2025-02-06 00:45:38 -08:00
  • ad3499858e clean moe align block kernel code and add acc test (#3332) Xiaoyu Zhang 2025-02-06 16:42:36 +08:00
  • 32de54ed1a [ROCm] Fix fp8 unrolledx4 matmul kernel. (#3325) Wen-Heng (Jack) Chung 2025-02-05 20:44:15 -06:00
  • 2d9c319594 Docker switch (#3327) saienduri 2025-02-05 18:06:50 -08:00
  • 07e58a2dcb update README (#3324) Yineng Zhang 2025-02-06 07:13:05 +08:00
  • 04d8cd2088 Initial Enablement of CI on MI300 (#3168) saienduri 2025-02-05 10:45:12 -08:00
  • a322051e31 Support custom mask for Triton attention (#3317) Ke Bao 2025-02-06 01:16:02 +08:00
  • de5533341e Update Triton extend backend interface (#3309) Ke Bao 2025-02-05 18:12:22 +08:00
  • 7aad8d1854 chore: bump v0.4.2.post2 (#3313) Yineng Zhang 2025-02-05 17:35:02 +08:00
  • 76fa2d152c Fix lora flashinfer import bug on ROCM (#3312) Baizhou Zhang 2025-02-05 00:36:49 -08:00
  • 7ab84948d8 [ROCm] Logic to decide whether to used manually unrolled kernel. (#3306) Wen-Heng (Jack) Chung 2025-02-04 21:12:20 -06:00
  • 4885b90802 Use forward_cuda to execute custom op for hip platform (#3305) kk 2025-02-05 10:58:17 +08:00
  • c2723a42a5 [ROCm] Manually unroll _w8a8_block_fp8_matmul kernel on AMD GPU. (#3299) Wen-Heng (Jack) Chung 2025-02-04 17:15:40 -06:00