Commit Graph

  • c7256ca836 [ROCm] Add tuning configs for AMD Radeon Graphics. (#3294) Wen-Heng (Jack) Chung 2025-02-04 12:34:57 -06:00
  • 6186a8f889 update flashinfer install index url (#3293) Yineng Zhang 2025-02-05 00:44:35 +08:00
  • a07364ccc5 Update Triton decode backend interface (#3292) Ke Bao 2025-02-04 23:26:04 +08:00
  • 2c1a695ff1 ROCm: sgl-kernel enablement starting with sgl_moe_align_block (#3287) HAI 2025-02-04 05:44:44 -08:00
  • d39899e85c upgrade flashinfer v0.2.0.post2 (#3288) Yineng Zhang 2025-02-04 21:41:40 +08:00
  • 70817a7eae [Feature] Define backends and add Triton backend for Lora (#3161) Baizhou Zhang 2025-02-03 22:09:13 -08:00
  • 7b5a374114 Update server args doc (#3273) simveit 2025-02-04 00:39:41 +01:00
  • 4b6f62e2bc add Atlas Cloud for Adoption and Sponsorship (#3276) Yineng Zhang 2025-02-04 05:31:30 +08:00
  • 897e2e253a add Nebius for Adoption and Sponsorship (#3274) Yineng Zhang 2025-02-04 04:41:26 +08:00
  • d54cee1441 adding Triton configs for DeepSeekV3 on Blackwell (#3272) kushanam 2025-02-03 12:12:09 -08:00
  • 00fa7d0417 add copyright for sgl-kernel (#3270) Yineng Zhang 2025-02-03 21:34:44 +08:00
  • 013021b6a1 refactor EAGLE 2 (#3269) Yineng Zhang 2025-02-03 20:52:30 +08:00
  • 3c8ac78dc1 optimize test_fused_moe style (#3268) Xiaoyu Zhang 2025-02-03 18:56:18 +08:00
  • 455bfe8dd3 Add a Doc about guide on nvidia jetson #3182 (#3205) Liangjun Song 2025-02-03 15:29:10 +11:00
  • 28b0a62bb3 Bug: Fix min_p sampling crash when using flashinfer backend (#3207) zifeitong 2025-02-02 15:36:07 -08:00
  • 566d61d90f ROCm: bump 6.3.0 (#3259) HAI 2025-02-02 12:13:40 -08:00
  • 55f5fc68ac Docs: Update accuracy evaluation (#3261) Chayenne 2025-02-02 11:14:59 -08:00
  • c27c378a19 docs/accuracy evaluation (#3114) simveit 2025-02-02 20:01:39 +01:00
  • d9eb9358cc Tune paged attention parameters for AMD GPU. (#3255) Wen-Heng (Jack) Chung 2025-02-01 19:29:45 -06:00
  • 959dca4fc7 use srt VocabParallelEmbedding (#3252) Yineng Zhang 2025-02-01 22:23:09 +08:00
  • f2b3a3188e Update README Yineng Zhang 2025-02-01 21:19:15 +08:00
  • ad6740977b add contact us in README (#3251) Yineng Zhang 2025-02-01 19:47:44 +08:00
  • 8db776f049 support QuickGELU (#3250) Yineng Zhang 2025-02-01 19:31:47 +08:00
  • 4eb4b401cc update and simplify CustomOp (#3249) Yineng Zhang 2025-02-01 18:56:44 +08:00
  • 17dbf976c5 update ENV to ROCm dockers (#3248) HAI 2025-02-01 01:27:43 -08:00
  • 5317902670 Add test for fp8 torch compile (#3246) Ke Bao 2025-02-01 16:07:54 +08:00
  • d7c0b32f4d [Docs] Add more details to profiling docs (#3221) Wenxuan Tan 2025-01-31 17:59:28 -06:00
  • 7b020cca2d add tuning block wise fp8 (#3242) Yineng Zhang 2025-02-01 03:58:18 +08:00
  • 7876279ea7 update cutlass dependency (#3240) Yineng Zhang 2025-02-01 03:13:44 +08:00
  • 34e405e01f update sgl-kernel version for sglang (#3238) Yineng Zhang 2025-02-01 02:14:41 +08:00
  • 1ebe1d6de5 Optimize MoE topk with torch compile (#3236) Ke Bao 2025-02-01 01:36:50 +08:00
  • 7811bfdaa7 compatible with flashinfer v0.2 (#3235) Yineng Zhang 2025-02-01 01:32:18 +08:00
  • 656f7fc1bc Docs: Quick fix for Speculative_decoding doc (#3228) Jhin 2025-01-31 10:30:40 -06:00
  • cf0f7eafe6 chore: bump v0.4.2.post1 (#3233) Yineng Zhang 2025-01-31 20:35:55 +08:00
  • b49d6d0fee support 12.5 CUDA runtime (#3231) Yineng Zhang 2025-01-31 20:31:38 +08:00
  • c02e313914 Fix block wise fp8 torch compile (#3232) Ke Bao 2025-01-31 19:56:02 +08:00
  • 734daedd8f [fix] Clamp logprob with dtype min to prevent -inf (#3224) Byron Hsu 2025-01-31 01:04:04 -08:00
  • 3ee62235c6 revert the MoE dependence (#3230) Yineng Zhang 2025-01-31 16:51:41 +08:00
  • 9829e77e3f Docs: Update supported models with Mistral 3 (#3229) Ravi Theja 2025-01-31 13:31:46 +05:30
  • cde4bbd5cc docs: add Novita for adoption and sponsorship (#3227) Ying Sheng 2025-01-30 18:28:22 -08:00
  • 9602c2aac7 keep the parts needed for moe_kernels (#3218) Yineng Zhang 2025-01-31 00:39:47 +08:00
  • e81d7f11de add tensorrt_llm moe_gemm as 3rdparty (#3217) Yineng Zhang 2025-01-30 23:49:14 +08:00
  • 222ce6f1da add tensorrt_llm common and cutlass_extensions as 3rdparty (#3216) Yineng Zhang 2025-01-30 23:04:41 +08:00
  • 468d23cff9 update setup for sgl-kernel (#3214) Yineng Zhang 2025-01-30 19:47:50 +08:00
  • c38b5fb4f4 update 3rdparty and rms norm for sgl-kernel (#3213) Yineng Zhang 2025-01-30 19:32:21 +08:00
  • 20453cef62 [test] Lower number of top logprobs to get rid of -inf (#3212) Byron Hsu 2025-01-30 02:01:23 -08:00
  • 9f635ea50d [Fix] Address remaining issues of supporting MiniCPMV (#2977) Mick 2025-01-28 16:22:13 +08:00
  • 76285fdeea Fix typo in README (#3190) Fidel González 2025-01-28 02:15:24 -05:00
  • 988d0a4bfc [kernel] Use sgl_kernel rope (#3169) Byron Hsu 2025-01-27 22:33:11 -08:00
  • 81262c7b72 clean up useless file (#3192) Xiaoyu Zhang 2025-01-28 14:29:30 +08:00
  • 27aeb4b7d8 [test] deduplicate test_session_control (#3183) Byron Hsu 2025-01-27 21:17:06 -08:00
  • 7b9b4f4426 Docs fix about EAGLE and streaming output (#3166) Jhin 2025-01-27 20:10:45 -06:00
  • 08104b56de Sanity check to prevent performance regression (#3171) Zhiqiang Xie 2025-01-27 12:28:17 -08:00
  • cf142b6eb8 fix: update Dockerfile for cu118 (#3181) Yineng Zhang 2025-01-27 23:46:44 +08:00
  • 4ab43cfb3e chore: bump v0.4.2 (#3180) Yineng Zhang 2025-01-27 21:42:05 +08:00
  • 2f79f58873 feat: use sgl-kernel 0.0.3 in sglang (#3179) Yineng Zhang 2025-01-27 21:39:52 +08:00
  • 8a96f74988 chore: bump 0.0.3 for sgl-kernel (#3178) Yineng Zhang 2025-01-27 20:29:28 +08:00
  • 827aa8730b cleanup sgl-kernel kernels (#3175) Yineng Zhang 2025-01-27 19:11:01 +08:00
  • f8ca66fb49 Update thresholds in test_nightly_gsm8k_eval.py (#3176) Lianmin Zheng 2025-01-27 03:02:09 -08:00
  • 53cef81587 Improve weight loading and code style (#3174) Lianmin Zheng 2025-01-27 03:00:41 -08:00
  • 351a72d40b add dsv3 mi300 triton config for block scale (#3146) yigex 2025-01-27 17:25:53 +08:00
  • 514f37c32b [kernel] Fix position ids in rope (#3173) Byron Hsu 2025-01-27 01:09:51 -08:00
  • 52c03f16b9 Add activation parameters to fused_moe (#3170) Lianmin Zheng 2025-01-27 00:23:37 -08:00
  • 741fccd7bf Bump sgl kernel to 0.0.2.post19 (#3167) Byron Hsu 2025-01-26 23:36:07 -08:00
  • 1e3e521544 add unit test for block wise fp8 (#3156) yizhang2077 2025-01-27 15:32:04 +08:00
  • fb11a43981 [kernel] Integrate flashinfer's rope with higher precision and better perf (#3134) Byron Hsu 2025-01-26 23:28:00 -08:00
  • af02f99b7c Add more logprob tests (#3162) Lianmin Zheng 2025-01-26 22:24:55 -08:00
  • 9472e69963 Doc: Add Docs about EAGLE speculative decoding (#3144) Jhin 2025-01-26 19:49:13 -06:00
  • 1acc1f561a [Docs]: Add function calling in index.rst (#3155) Chayenne 2025-01-26 11:11:27 -08:00
  • b045841bae Feature/function calling update (#2700) YAMY 2025-01-26 09:57:51 -08:00
  • f265d15b96 use self-hosted to build sgl-kernel (#3154) Yineng Zhang 2025-01-26 23:02:57 +08:00
  • 02431b9ad2 fix link in README (#3153) Yineng Zhang 2025-01-26 21:30:00 +08:00
  • 1dda8c5e4c Return more infos for computing average acceptance length (#3152) Lianmin Zheng 2025-01-26 04:51:54 -08:00
  • 7e0976133c udpate sgl-kernel version for srt (#3150) Yineng Zhang 2025-01-26 20:22:34 +08:00
  • f4a92f4b56 Temporarily skip the openai frontend tests (#3151) Lianmin Zheng 2025-01-26 04:17:35 -08:00
  • 318260c0fa chore: bump 0.0.2.post18 for sgl-kernel (#3149) Yineng Zhang 2025-01-26 19:00:34 +08:00
  • 4a61253123 Do not load OPENAI_KEY from secrets (#3147) Lianmin Zheng 2025-01-26 01:54:03 -08:00
  • d1a0863251 Add a test case for cached_tokens (#3145) Lianmin Zheng 2025-01-26 01:39:28 -08:00
  • f8b28e461a Add CPU affinity setting to latency benchmark (#3085) Hubert Lu 2025-01-25 23:52:05 -08:00
  • 82392da830 support w8a8 fp8 kernel with CUTLASS (#3047) HandH1998 2025-01-26 15:46:51 +08:00
  • 95f789adb0 minor: cleanup sgl-kernel (#3143) Yineng Zhang 2025-01-26 14:29:58 +08:00
  • 4f118a39d7 Fix repetition penalty (#3139) Lianmin Zheng 2025-01-25 21:48:58 -08:00
  • 66283dbc0c [Fix] Not skip NVML Check on AMD Platform (#3135) yigex 2025-01-26 13:33:51 +08:00
  • 822bae8c00 feat: cross python wheel for sgl-kernel (#3138) Yineng Zhang 2025-01-26 13:21:34 +08:00
  • 8e48ca8cc1 enable kv_scale for Gemma2 (#3113) Hui Liu 2025-01-25 18:29:14 -08:00
  • 27acf63bbd Use torch.compile for scaling penalty (#3133) Lianmin Zheng 2025-01-25 18:27:33 -08:00
  • da6f8081f6 Fix CI tests (#3132) Lianmin Zheng 2025-01-25 17:43:39 -08:00
  • 9286740eff feat: refactor sgl-kernel and use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#3130) yinfan98 2025-01-26 02:55:08 +08:00
  • 896c07441e update installation doc for sgl-kernel (#3129) Yineng Zhang 2025-01-26 00:00:13 +08:00
  • c23d5706f4 Update whl index path (#3128) Ke Bao 2025-01-25 23:57:09 +08:00
  • 67ad4338e1 Update tag name for whl release (#3127) Ke Bao 2025-01-25 23:14:35 +08:00
  • 3cab5f71ea speedup pr test for sgl-kernel (#3126) Yineng Zhang 2025-01-25 21:37:48 +08:00
  • 14e754a868 chore: bump v0.0.2.post17 for sgl-kernel (#3125) Yineng Zhang 2025-01-25 20:43:02 +08:00
  • 98522149ff mirror fix for custom allreduce (#3124) yizhang2077 2025-01-25 18:26:41 +08:00
  • 5d9d15e70f support fp32 in sampling_scaling_penalties kernel (#3121) Xiaoyu Zhang 2025-01-25 16:52:17 +08:00
  • 665e5e85f6 Add step to update sgl-kernel whl index (#3110) Ke Bao 2025-01-25 02:03:01 +08:00
  • a22f60a313 Add workflow for sgl-kernel cu118 release (#3109) Ke Bao 2025-01-24 22:30:30 +08:00
  • 04f0b4cbef minor: update sgl-kernel setup (#3107) Yineng Zhang 2025-01-24 20:10:35 +08:00
  • 4505a43614 [Docs] minor update for phi-3 and phi-4 (#3096) Adarsh Shirawalmath 2025-01-24 17:30:20 +05:30
  • 685a5738a7 Allow local cutlass directory to be used in sgl-kernel build (#3037) Trevor Morris 2025-01-24 03:59:47 -08:00