Commit Graph

  • 7130a7cea9 refine sgl_moe_align_block_size_benchmark (#4327) Xiaoyu Zhang 2025-03-12 13:48:38 +08:00
  • 8f1f614ee2 [Docs] Clean up benchmark_and_profiling.md (#4297) Michael Yao 2025-03-12 12:48:21 +08:00
  • 7140ba3573 Add A800 tuning configs for DeepSeek R1/V3 channel-wise INT8 (#4323) lambert0312 2025-03-12 09:25:56 +08:00
  • d1da58e275 unify is_cuda and is_hip (#4321) Yineng Zhang 2025-03-11 18:12:56 -07:00
  • 1cf63485c1 upgrade flashinfer 0.2.3 (#4317) Yineng Zhang 2025-03-11 15:37:17 -07:00
  • ff2ce0b86f refactor: move image processors to separate files (#4229) Mick 2025-03-12 03:35:35 +08:00
  • 0f2a2e3c19 Add H20 tuning configs support DeepSeek V3/R1 INT8(block-wise) (#4220) Ximingwang-09 2025-03-12 03:32:33 +08:00
  • 690e1f2371 [AMD] Fix rocm sgl-kernel missing modules error (#4311) yigex 2025-03-12 01:35:28 +08:00
  • 00f42707ea update doc (#4299) Yineng Zhang 2025-03-11 01:14:16 -07:00
  • 6a02b32d07 Add A100 tuning configs for DeepSeek R1/V3 channel-wise INT8 (#4287) yych0745 2025-03-11 15:49:06 +08:00
  • 3a08f54638 Update MTP doc (#4290) Ke Bao 2025-03-11 15:46:55 +08:00
  • dce303e279 linear support deepgemm (#4199) lukec 2025-03-11 15:38:37 +08:00
  • 4d27eb9ad1 update sgl-kernel 0.0.4.post2 (#4291) Yineng Zhang 2025-03-11 00:34:33 -07:00
  • d3ecd63204 Add A800 tuning configs support DeepSeek V3/R1 BF16 and INT8(block-wise) (#4136) lambert0312 2025-03-11 15:32:25 +08:00
  • cd90945518 bump sgl-kernel 0.0.4.post2 (#4288) Yineng Zhang 2025-03-11 00:09:47 -07:00
  • bde24ab31f update deepgemm (#4284) Yineng Zhang 2025-03-10 23:39:57 -07:00
  • bf2eefc0c7 Uupdate cutalss dependency for its bug fix (#4277) Elfie Guo 2025-03-10 17:00:05 -07:00
  • 5524e7d057 Fix nightly eval for neuralmagic/Mixtral-8x7B-Instruct-v0.1-FP8 (#4279) Lianmin Zheng 2025-03-10 16:50:28 -07:00
  • e187a3d595 upgrade xgrammar 0.1.15 (#4275) Yineng Zhang 2025-03-10 14:53:24 -07:00
  • 3dd4feae63 add THIRDPARTYNOTICES for DeepGEMM (#4272) Yineng Zhang 2025-03-10 11:10:57 -07:00
  • 2ac189edc8 Amd test fp8 (#4261) HandH1998 2025-03-11 01:12:09 +08:00
  • 5a6400eec5 Test no vllm custom allreduce (#4256) Lianmin Zheng 2025-03-10 10:08:25 -07:00
  • cf0ccd406e Optimize rope in sgl kernel (#4267) Lianmin Zheng 2025-03-10 10:07:45 -07:00
  • 3d56585a97 increase the timeout of nightly-test.yml (#4262) Lianmin Zheng 2025-03-10 05:07:03 -07:00
  • 00d25a7f5e Fix quantization and nightly tests (#4258) Lianmin Zheng 2025-03-10 03:06:21 -07:00
  • 1a5023e05d Release sgl-kernel v0.0.4.post1 (#4255) Lianmin Zheng 2025-03-10 02:39:50 -07:00
  • 23308a9032 fix per_token_group_quant_fp8 illegal memory when num_groups % 16 != 0 (#4231) Xiaoyu Zhang 2025-03-10 16:42:58 +08:00
  • ac69885056 fix the input_ids is None error (#4144) shimin 2025-03-10 16:38:37 +08:00
  • aa957102a9 Simplify tests & Fix trtllm custom allreduce registration (#4252) Lianmin Zheng 2025-03-10 01:24:22 -07:00
  • 007f8b3dc2 Added example for multimodal embedding (#4206) simveit 2025-03-10 08:53:56 +01:00
  • 4455b26e76 [Bug fixed] fixed the crash when enable the dp-attention on the single card (#3958) DavidChan 2025-03-10 15:50:34 +08:00
  • c553e1604c DeepGemm integrate to sgl-kernel (#4165) laixin 2025-03-10 15:35:07 +08:00
  • 7c0541b385 Move activation.cu to sgl-kernel/elementwise (#4250) Lianmin Zheng 2025-03-09 22:41:13 -07:00
  • e8a69e4d0c Clean up fp8 support (#4230) Lianmin Zheng 2025-03-09 21:46:35 -07:00
  • fbd560028a Auto balance CI tests (#4238) Lianmin Zheng 2025-03-09 21:05:55 -07:00
  • 730d084f2a Minor style fix for sgl-kernel (#4243) Lianmin Zheng 2025-03-09 20:15:13 -07:00
  • 4a05bdfa86 Revert "Check eagle server args" (#4242) Lianmin Zheng 2025-03-09 18:53:33 -07:00
  • eb06dbcbf8 Move rope and bmm into sgl-kernel (#4241) Lianmin Zheng 2025-03-09 18:38:15 -07:00
  • 9dfafa743c Fix test of flashinfer mla with nextn (#4237) Baizhou Zhang 2025-03-09 12:45:39 -07:00
  • f1d09a6541 Update bench speculative script (#4235) Ke Bao 2025-03-10 03:19:01 +08:00
  • df84ab2a5b update sgl-kernel 3rdparty (#4228) Yineng Zhang 2025-03-09 01:16:05 -08:00
  • 34c8898755 Check eagle server args (#4217) Ying Sheng 2025-03-09 01:10:43 -08:00
  • 0dd6cda288 Apply sgl w8a8 fp8 kernel (#3148) HandH1998 2025-03-09 16:03:32 +08:00
  • 9fb48f951f Support nextn for flashinfer mla attention backend (#4218) Baizhou Zhang 2025-03-09 00:01:54 -08:00
  • 89ccb533ad use sgl-kernel 0.0.4 (#4224) Yineng Zhang 2025-03-08 23:43:09 -08:00
  • dceb256f1b [docs] Unhide production metrics page (#4193) Stefan He 2025-03-08 23:41:40 -08:00
  • 0e90ae628a [docker] Distributed Serving with k8s Statefulset ( good example for DeepSeek-R1) (#3631) Peter Pan 2025-03-09 15:41:20 +08:00
  • 1361ab9e03 Lazily import lora backends (#4225) Lianmin Zheng 2025-03-08 23:39:26 -08:00
  • 5c7dd14ba1 chore: bump v0.0.4 for sgl-kernel (#4223) Yineng Zhang 2025-03-08 23:01:59 -08:00
  • 8abf74e3c9 Rename files in sgl kernel to avoid nested folder structure (#4213) Lianmin Zheng 2025-03-08 22:54:51 -08:00
  • ee132a4515 use latest sgl-kernel for mla test (#4222) Yineng Zhang 2025-03-08 22:27:47 -08:00
  • 79a321af55 revert pr 3628 to pass test_mla ci (#4219) Xiaoyu Zhang 2025-03-09 13:15:14 +08:00
  • 6eec3cdce6 docs(reasoning content): 📝 deepseek-r1 parser support qwq (#4124) Xihuai Wang 2025-03-09 12:14:50 +08:00
  • 48473684cc Split test_mla.py into two files (#4216) Lianmin Zheng 2025-03-08 15:40:49 -08:00
  • b3251e9f40 refine quant kernel code style (#4211) Xiaoyu Zhang 2025-03-08 21:47:35 +08:00
  • 2cadd51d11 Test no vllm custom allreduce (#4210) Lianmin Zheng 2025-03-08 05:23:06 -08:00
  • 4a893d142d Refactor Dockerfile: unify CUDA logic and reduce image size by ~2.6 GB (#3749) Kebe 2025-03-08 19:01:13 +08:00
  • 8d323e95e4 Use clang format 18 in pr-test-sgl-kernel.yml (#4203) Lianmin Zheng 2025-03-08 01:28:10 -08:00
  • 0fe7c13be1 Fix bench_serving flush cache not recognizing OPENAI_API_KEY (#4181) Mingshan 2025-03-08 17:03:38 +08:00
  • 08c4d764a5 lazy import attn backends (#4200) Lianmin Zheng 2025-03-08 00:41:35 -08:00
  • 96d0e37fa7 Revert "Minor improvement to per_tensor_quant_fp8 (#4197)" (#4198) Yineng Zhang 2025-03-07 22:57:09 -08:00
  • 90bb2be27e Minor improvement to per_tensor_quant_fp8 (#4197) Rex 2025-03-07 22:52:12 -08:00
  • b93ef5e56d Remove the vllm dependency from the moe_align function (#4164) lukec 2025-03-08 14:42:16 +08:00
  • d4017a6b63 [EAGLE] many fixes for eagle (#4195) Lianmin Zheng 2025-03-07 22:12:13 -08:00
  • d052f4c8a9 New clang format for sgl kernel (#4194) Lianmin Zheng 2025-03-07 20:21:08 -08:00
  • e1aaa79ac9 Update amd ci docker image to v0.4.3.post4-rocm630. (#4189) saienduri 2025-03-07 13:02:02 -08:00
  • 20c8119915 Fix eagle hang issue for max_new_tokens=1 (#4185) Ke Bao 2025-03-08 04:11:18 +08:00
  • 70866b6f4f use same version for ci and pyproject (#4187) Yineng Zhang 2025-03-07 10:39:55 -08:00
  • eb61f5c9af Revert "ROCm: Flex Attention Enablement with custom backends (#4178)" (#4186) Yineng Zhang 2025-03-07 10:27:52 -08:00
  • 0beea4503f ROCm: Flex Attention Enablement with custom backends (#4178) HAI 2025-03-07 04:38:53 -08:00
  • c827c671f7 [Docs] Improve bullets appearance and grammar (#4174) Michael Yao 2025-03-07 19:16:25 +08:00
  • b55a621ffb fix int8 doc link (#4179) Yineng Zhang 2025-03-07 02:49:19 -08:00
  • ffa1b3e318 Add an example of using deepseekv3 int8 sglang. (#4177) lukec 2025-03-07 17:56:09 +08:00
  • 7e3bb52705 update release-pypi-kernel Yineng Zhang 2025-03-07 01:48:47 -08:00
  • 96263f275c chore: bump v0.0.3.post7 for sgl-kernel (#4176) Yineng Zhang 2025-03-07 01:15:34 -08:00
  • 9376ac361d Memory pool fix for upstream change about eagle (#4170) Zhiqiang Xie 2025-03-07 00:58:20 -08:00
  • 94a2b9d33e Put utils in ifndef USE_ROCM to fix CI (#4167) (#4168) Yineng Zhang 2025-03-07 00:01:17 -08:00
  • 3c3eb374b2 Remove non-existent AMD header include (#4166) Stefan He 2025-03-06 23:29:30 -08:00
  • d557319a8b [Docs] Fix links and grammar issues (#4162) Michael Yao 2025-03-07 15:14:18 +08:00
  • 95085d65e9 [Refactor] Reducing code duplication across FP8 CUDA quantization kernels (#4163) Stefan He 2025-03-06 22:58:52 -08:00
  • c7f254468f [Feature] DeepSeek V3/R1 INT8 Quantization (channel-wise) (#3888) HandH1998 2025-03-07 12:54:52 +08:00
  • 63ee26d162 Add sgl_per_token_quant_fp8 (#4089) Stefan He 2025-03-06 20:53:05 -08:00
  • ad55f17182 [quant kernel] sgl-kernel support per_tensor_quant fp8 (#3786) Xiaoyu Zhang 2025-03-07 10:05:43 +08:00
  • 361971b859 Add Support for Qwen2-VL Multi-modal Embedding Models (#3694) Pan Lyu 2025-03-07 08:46:20 +08:00
  • 13bc39c5d6 ROCm: enable trillion-parameter MoE models with INT4-FP8 single node (#4152) HAI 2025-03-06 15:33:02 -08:00
  • 9854a18a51 Hot fix small vocal eagle in docs (#4154) Chayenne 2025-03-06 15:13:26 -08:00
  • ebddb65aed Docs: add torch compile cache (#4151) Chayenne 2025-03-06 14:27:09 -08:00
  • 19fd57bcd7 [docs] fix HF reference script command (#4148) Adarsh Shirawalmath 2025-03-07 02:51:54 +05:30
  • 9c58e68b4c Release v0.4.3.post4 (#4140) Lianmin Zheng 2025-03-06 12:50:28 -08:00
  • d03b3467b8 Fix constrained generation errors by adding datasets dependency (#4142) Oliver Stanley 2025-03-06 20:07:51 +00:00
  • ab7fba0ece Fix nightly ci Gsm8k & Fix flashinfer backend kvcache quant (#4147) yinfan98 2025-03-07 03:50:07 +08:00
  • bc1534ff32 Fix a draft model accuracy bug in eagle; support step=1; return logprob in eagle (#4134) Lianmin Zheng 2025-03-06 06:13:59 -08:00
  • 3a3918121f fix bench serving bug (#4135) Lzhang-hub 2025-03-06 21:34:02 +08:00
  • 800bf018fb Update CODEOWNER (#4138) Lianmin Zheng 2025-03-06 03:42:10 -08:00
  • b16af90bc3 AMD/ROCm: update base image string (#4137) kk 2025-03-06 19:38:54 +08:00
  • 98c73d71cb [Minor] make the __init__ function of model_runner.py shorter (#4132) Lianmin Zheng 2025-03-06 01:51:12 -08:00
  • fcc2e37f69 Split the __init__ of scheduler as smaller functions. Improve the eagle tests (#4128) Lianmin Zheng 2025-03-06 00:13:20 -08:00
  • 0804dd11a0 remove unused max_jobs in setup_rocm.py (#4126) Liu Jinjie 2025-03-06 00:12:19 -08:00
  • 55dc8e4d52 Add tag suffix to nightly docker builds. (#4129) saienduri 2025-03-05 23:22:36 -08:00
  • 02e9e9f1cf Add codeowners for eagle implementations (#4131) Ying Sheng 2025-03-05 23:16:49 -08:00