Commit Graph

  • 23c764b18a [Feature] Support DeepEP Low Latency (#4767) Jinyan Chen 2025-04-02 00:23:25 +08:00
  • 87fafa0105 Revert PR 4764 & 4813 related to R1 RoPE (#4959) Yuhong Guo 2025-04-01 11:56:58 +08:00
  • 1c63e79756 use fa3 in sgl-kernel (#4954) Yineng Zhang 2025-03-31 16:14:49 -07:00
  • ee47a6c1c3 [Build] Fix cuda12.8 build error in nvfp4_scaled_mm_kernels.cu (#4953) Yuhong Guo 2025-04-01 03:00:34 +08:00
  • 6384d31776 bump sgl-kernel v0.0.6 (#4950) Yineng Zhang 2025-03-31 11:24:09 -07:00
  • 5cb552b1d4 refactor: multimodal data (#4754) Mick 2025-04-01 00:57:51 +08:00
  • c7457191a0 [Fix] revert clean m.def for cudagraph (#4944) yinfan98 2025-03-31 17:08:55 +08:00
  • 51ac297ace [feat] interface for platforms abstraction (#4928) JieXin Liang 2025-03-31 15:04:21 +08:00
  • a169b9f813 Fix oom error for large page size (#4913) Zhiqiang Xie 2025-03-30 21:34:21 -07:00
  • 4a63bc32b7 [Fix] Add torch compile for torch.clamp back (#4936) Baizhou Zhang 2025-03-30 20:46:07 -07:00
  • a303325fdb Fix DeepSeek bug causing 2.2% MMLU drop when TP!=DP (#4883) fzyzcjy 2025-03-31 11:10:21 +08:00
  • 42873eac09 [Fix] Improve Lora tests and reduce CI runtime (#4925) Baizhou Zhang 2025-03-30 19:40:14 -07:00
  • 4814ecaff9 cleanup sgl-kernel (#4933) Yineng Zhang 2025-03-30 14:12:30 -07:00
  • e62d60fe6d [Fix] avoid stream sync and torch compile in prefill for fa3 backend (#4932) Baizhou Zhang 2025-03-30 13:53:44 -07:00
  • 032f8faaab Fix sglang frontend's incorrect dependency on torch (#4931) SEPLOS 2025-03-31 04:00:24 +08:00
  • 37c66ec856 [feat] add fa3 in sgl-kernel (#4902) yinfan98 2025-03-31 03:57:10 +08:00
  • 9adf178cc2 Fix 2-gpu CI test and suppress some warnings (#4930) Lianmin Zheng 2025-03-30 12:51:44 -07:00
  • f842853a40 Fix the timeout for unit-test-2-gpu in pr-test.yml (#4927) Lianmin Zheng 2025-03-30 12:15:40 -07:00
  • 195a09f57c fix bmm fp8 (#4926) Yineng Zhang 2025-03-30 12:15:20 -07:00
  • 9fccda3111 [Feature] use pytest for sgl-kernel (#4896) Adarsh Shirawalmath 2025-03-30 23:06:52 +05:30
  • 4ede6770cd Fix retract for page size > 1 (#4914) Lianmin Zheng 2025-03-30 02:57:15 -07:00
  • b26bc86b36 Support page size > 1 + eagle (#4908) Lianmin Zheng 2025-03-30 00:46:23 -07:00
  • 5ec5eaf760 fix allreduce test (#4909) Yi Zhang 2025-03-30 14:16:53 +08:00
  • 0d7fe866f9 [Misc] Clean m.def and add Development Tips (#4890) yinfan98 2025-03-30 14:06:18 +08:00
  • 54b9a2de0a remove setup for sgl-kernel (#4899) Yineng Zhang 2025-03-29 12:47:38 -07:00
  • 8e7b31546c quick fix: add default for new kernel (#4898) yinfan98 2025-03-30 03:31:59 +08:00
  • 45dcfc2e76 Add deepseek style fused moe group gate selection kernel (#4530) Qingquan Song 2025-03-29 11:51:45 -07:00
  • ddf8981d91 Delete test_deep_gemm.py (#4891) yinfan98 2025-03-30 01:46:11 +08:00
  • 400ad66019 Update CODEOWNERS (#4889) Yineng Zhang 2025-03-29 09:56:51 -07:00
  • 05625b9792 [Docs] Update DeepGEMM at README.md (#4886) yinfan98 2025-03-30 00:53:39 +08:00
  • 736502d4fd Tiny fix doc error (#4795) fzyzcjy 2025-03-29 23:22:17 +08:00
  • 8690c40bb0 Improve stack trace of retry errors (#4845) fzyzcjy 2025-03-29 23:21:31 +08:00
  • b1cfb4e972 Fix BadRequestError wrong arguments and remove openai dependency (#4882) fzyzcjy 2025-03-29 23:16:21 +08:00
  • 19e96e5923 bump v0.4.4.post3 (#4878) Yineng Zhang 2025-03-28 23:21:24 -07:00
  • aa08aeacf4 update torch compile doc (#4874) Ke Bao 2025-03-29 10:49:30 +08:00
  • d8a136a113 upgrade sgl-kernel 0.0.5.post4 (#4873) Yineng Zhang 2025-03-28 19:48:56 -07:00
  • 20c90be23d [Feature] Support FA3 backend for MLA (#4831) Baizhou Zhang 2025-03-28 18:30:14 -07:00
  • ec3ee0289d fix sgl-kernel cu118 build (#4872) Yineng Zhang 2025-03-28 17:23:51 -07:00
  • 92941ce7b5 bump sgl-kernel 0.0.5.post4 (#4768) Yineng Zhang 2025-03-28 14:40:53 -07:00
  • 2bb0e7cf43 fix sampling issue (#4871) Yineng Zhang 2025-03-28 14:07:21 -07:00
  • 72549263c6 update sgl-kernel test ci (#4866) Yineng Zhang 2025-03-28 11:42:41 -07:00
  • 044c315970 Make torch compile configurable for biased_grouped_topk (#4749) Qingquan Song 2025-03-28 10:57:52 -07:00
  • 4db29e82ec [Feat] support deepgemm for cmake (#4864) yinfan98 2025-03-29 01:51:44 +08:00
  • c483377ed7 Fix wrong variable name when stopping memory profile (#4772) Fr4nk1in 2025-03-29 01:35:02 +08:00
  • 74e0ac1dbd Clean up import vllm in quantization/__init__.py (#4834) Lianmin Zheng 2025-03-28 10:34:10 -07:00
  • ef9a378a20 [Feature] add multi-rank support for Lora (#4492) chaobo jia 2025-03-29 00:38:44 +08:00
  • 6dea5c96bf Revert "get the python version from env (#4729)" (#4863) Yineng Zhang 2025-03-28 08:07:48 -07:00
  • 6ffb6bd47a Fix fa3 cuda graph page_size > 1 precision and page_size=1 speed (#4855) Qingquan Song 2025-03-28 01:35:59 -07:00
  • 47e6628aae Fix CI tests (#4853) Lianmin Zheng 2025-03-28 00:28:35 -07:00
  • 7907f9eb20 test: reduce mem_fraction_static for gemma3 vision test (#4840) Juwan Yoo 2025-03-27 23:20:10 -07:00
  • 8c04f0f2e1 Support with_stack and record_shapes in profiler (#4740) fzyzcjy 2025-03-28 14:01:42 +08:00
  • 265e756494 Super tiny remove unused code (#4750) fzyzcjy 2025-03-28 13:32:14 +08:00
  • d3f71f5e19 Fix torch.cuda.MemPool() internal assertion failure (#4687) fzyzcjy 2025-03-28 13:29:36 +08:00
  • 5eae67cb1f get the python version from env (#4729) DavidChan 2025-03-28 13:26:42 +08:00
  • 6dbf99982f Fix missing arguments in SchedulePolicy and RadixCache initialization in tests. (#4712) vikram singh shekhawat 2025-03-28 10:53:51 +05:30
  • e0166f8ab4 Remove empty tool function name (#4704) Kebe 2025-03-28 13:23:30 +08:00
  • 53a2c3b466 Support controlling nsys start and end range programmatically (#4688) fzyzcjy 2025-03-28 13:21:13 +08:00
  • 550586ef42 fix: Inappropriate lack of Optional type on OpenAI ChatCompletionRequest (#4681) BroadbentJim 2025-03-28 05:19:05 +00:00
  • cf29fe9e78 Fix Engine error when enabling DP attention (#4648) fzyzcjy 2025-03-28 13:17:30 +08:00
  • 26c0f13126 Support Page Size > 1 for FA3 (#4832) Stefan He 2025-03-27 22:07:14 -07:00
  • f9970bd1af fix: when use SGLANG_PORT this env,port is str (#4528) rongfu.leng 2025-03-28 12:46:06 +08:00
  • 2e0f94ab79 [Fix] fix output_top_logprobs is not exist (#4597) lambert0312 2025-03-28 12:45:57 +08:00
  • 18317ddc13 ci: add condition for daily docker build (#4487) warjiang 2025-03-28 12:44:37 +08:00
  • e2e2ab70e0 IPv6 support (#3949) Vincent 2025-03-28 00:42:13 -04:00
  • 0d3e3072ee Fix CI of test_patch_torch (#4844) fzyzcjy 2025-03-28 12:22:45 +08:00
  • 62dd95870c Remove retry in nightly tests (#4846) fzyzcjy 2025-03-28 12:18:29 +08:00
  • 72031173e4 fix: fix typo of comments in w8a8_fp8.py (#4843) Jiaqi 2025-03-28 12:06:47 +08:00
  • 9fdc6d6abc Fix the lora adapter when lora path is none (#4799) Qiaolin Yu 2025-03-28 00:03:08 -04:00
  • 42a45df043 [Fix] self.worker assignment in TpModelWorker and refactor references (#4788) XinyuanTong 2025-03-27 20:28:38 -07:00
  • 04eb6062e4 Include context length in /v1/models response. (#4809) Jon Durbin 2025-03-27 23:23:18 -04:00
  • e84f4ba0ab [Misc] Fix issues reported by torchfix (#4837) Brayden Zhong 2025-03-27 23:10:32 -04:00
  • b149b39353 [CI] Remove unused imports with Ruff to pre-commit config, only to benchmarks/docs/examples folder (#3969) Brayden Zhong 2025-03-27 22:45:02 -04:00
  • 31dfff7da7 use default for torch.ops (#4835) Yineng Zhang 2025-03-27 19:09:58 -07:00
  • 10a9ab7b07 Fix error due to CustomAllreduce setup failure (#4815) Kebe 2025-03-28 09:52:10 +08:00
  • bb0fd749a6 [Fix] Add compressed_tensors as deps (#4819) Junrong Lin 2025-03-28 09:08:24 +08:00
  • 7f19e083c1 Support (1 <= dp < tp) in the dp attention in DeepEP (#4770) tarinkk 2025-03-27 20:09:35 -04:00
  • 98a2cfa9b2 Basic Cleanup (#4833) Daniel Holanda 2025-03-27 16:55:48 -07:00
  • 2a882e8f3a Fix the nightly eval by lowering the threshold of neuralmagic/gemma-2-2b-it-FP8 (#4830) Lianmin Zheng 2025-03-27 16:09:49 -07:00
  • e6e4d02245 Update MMMU Benchmark instructions (#4694) Ravi Theja 2025-03-28 03:14:16 +05:30
  • 188105a21b deps: lazy import optional dependencies gguf and torchvision (#4826) Juwan Yoo 2025-03-27 14:35:36 -07:00
  • b39532587b Update doc for DeepSeek-V3-0324 (#4825) Ke Bao 2025-03-28 04:30:40 +08:00
  • 5fa3058f01 fix the release doc dependency issue (#4828) Yineng Zhang 2025-03-27 13:28:12 -07:00
  • bbab97a6a8 add partial_json_parser and einops (#4827) Yineng Zhang 2025-03-27 13:24:54 -07:00
  • 0bc0bf5734 gemma3: impl get_attention_sliding_window_size for attn init (#4823) Juwan Yoo 2025-03-27 10:43:58 -07:00
  • f60f293195 [k8s] Clarified the usage of shared memory. (#4341) Jiří Suchomel 2025-03-27 16:53:19 +01:00
  • 17000d2b3a Remove Unintended Capture Batch Sizes in AMD HIP Graph Runner (#4638) AinL 2025-03-28 00:41:33 +09:00
  • 668ecc6c5b Fix ut mla-test-1-gpu-amd (#4813) strgrb 2025-03-27 23:27:51 +08:00
  • 886fcbdd09 Use apply_rope_with_cos_sin_cache_inplace for DeepSeek (#4764) strgrb 2025-03-27 16:45:37 +08:00
  • 8bf6d7f406 support cmake for sgl-kernel (#4706) Yineng Zhang 2025-03-27 01:42:28 -07:00
  • 1b9175cb23 [FA3 Attn Backend] Remove Unnecessary Device Sync for FA3 (#4745) Stefan He 2025-03-27 00:45:11 -07:00
  • 92bb49a7f9 Patch PyTorch's bug that cross-process tensor transfer will lead to wrong device (#4565) fzyzcjy 2025-03-27 15:22:33 +08:00
  • 6f5cc5eb05 update xgrammar 0.1.17 (#4804) Yineng Zhang 2025-03-27 00:21:59 -07:00
  • c913ed4046 support clip embedding model (#4506) Pan Lyu 2025-03-27 15:18:15 +08:00
  • 1afe3d0798 Align finish reason and stream mode in openai api (#4388) Xihuai Wang 2025-03-27 15:16:52 +08:00
  • 44f47d3ee1 Update supported_models.md: adding open-r1 Olympic Code 32B by HuggingFace (#4628) Didier Durand 2025-03-27 08:16:16 +01:00
  • ae25d36dc6 [3/3] fix dsv3 awq issue (#4719) laixin 2025-03-27 14:13:43 +08:00
  • 1099f6c974 bump v0.4.4.post2 (#4669) Yineng Zhang 2025-03-26 19:58:00 -07:00
  • 04e3ff6975 Support compressed tensors fp8w8a8 (#4743) Xiaoyu Zhang 2025-03-27 04:21:25 +08:00
  • 45fdf1f7f3 Fix shared memory OOM on sm86 GPUs. (#4797) Yi Pan 2025-03-27 01:41:53 +08:00
  • d89c0e4b7e Use metadata to detect version of package (#4782) Kebe 2025-03-26 15:41:43 +08:00