Commit Graph

  • 432f2053dd [sgl-kernel] 1/N Refactor sglang cutlass 3x - gemm fp8 blockwise sm90 (#8913) Yuan Luo 2025-08-15 01:55:54 +08:00
  • 1fea998a45 chore: bump sgl-kernel v0.3.5 (#9185) Yineng Zhang 2025-08-14 03:20:48 -07:00
  • 5aa1ebd242 [2/n]decouple quantization implementation from vLLM dependency (#8112) Peng Zhang 2025-08-14 18:19:03 +08:00
  • 4dbf43601d fix: zero_init buffer (#9065) eigen 2025-08-14 05:39:09 -04:00
  • 3d6be1fbce add w8a8-fp8-block-wise H20-3e triton config (#8018) lukec 2025-08-14 14:15:09 +08:00
  • 4063234c1a Add H200 fused MoE kernel configs for DeepSeek-V3 in triton 3.3.1 (#7687) Jun Liu 2025-08-14 15:14:09 +09:00
  • 83feef5b2c Add H20 fused MoE kernel configs for Dpsk & Qwen3 (#7631) Tommy Yang 2025-08-14 14:13:22 +08:00
  • 2871eacc05 Add Triton Fused MoE kernel config for E=16 on B200 (#7004) Brayden Zhong 2025-08-14 02:12:27 -04:00
  • ac15bdc194 Add H200 fused MoE kernel tuning configs for Qwen3-Coder-480B-A35B-Instruct (#8852) forestlee95 2025-08-14 14:11:11 +08:00
  • d6451c3f65 Add A800 fused MoE kernel tuning configs for GLM4.5 and GLM4.5-Air (#8808) Li Hui 2025-08-14 14:03:17 +08:00
  • 841810f227 [Perf] Tunings for SM100 FP8 CUTLASS kernel (#8818) henryg 2025-08-13 21:59:22 -07:00
  • 733446dd36 fix io group (#9154) pansicheng 2025-08-14 12:46:42 +08:00
  • 4c22897a66 Feature: support qwen and llama4 reducescatter for dp attention padding (#9101) wxzhoucs 2025-08-14 12:10:29 +08:00
  • 1bc183c6de Faster weight processing (trtllm-gen moe nvfp4) (#9162) Alex Yang 2025-08-13 21:09:34 -07:00
  • b87aacb5c5 [DP Attention] Refactor: adding some utility functions (#9136) Cheng Wan 2025-08-13 21:08:06 -07:00
  • b3363cc1aa Fix docker container DeepEP error on Blackwell (#9171) fzyzcjy 2025-08-14 12:06:48 +08:00
  • 98457c0453 [Bugfix] Avoid unnecessary reduce-scatter call in prepare_mlp (#9169) Huaixin Chang 2025-08-14 12:04:41 +08:00
  • 0fc8bf2cd4 [AMD] Update fallback images for AMD CI (#9159) michael-amd 2025-08-13 20:15:10 -07:00
  • a669bc2f74 Replace sglang.srt.layers.quantization.scalar_types with sgl_kernel.scalar_type (#8951) Hongbo Xu 2025-08-14 10:41:41 +08:00
  • 6b7c24712c Fix broken trtllm_mha attn backend with gpt-oss (#9161) Nicolas Castet 2025-08-13 18:40:55 -05:00
  • a027a9b4b3 [Generative Score API] Optimization to Remove Decode. (#8840) Sundara Raman Ramachandran 2025-08-13 14:12:24 -07:00
  • 9e426466af Clean up allocators (#9134) Lianmin Zheng 2025-08-13 13:56:04 -07:00
  • 2f20f43026 Swap xeon ci to gnr server (#9042) DiweiSun 2025-08-14 03:39:19 +08:00
  • 65736dc524 [Model] Support Qwen3ForSequenceClassification for Qwen3-Embed Model (#7957) Zhihao Liu 2025-08-14 02:14:54 +08:00
  • 7b56e494be chore: bump v0.5.0rc1 (#9069) Yineng Zhang 2025-08-13 10:44:14 -07:00
  • 0ff6d1fce1 Support FA3 backend for gpt-oss (#9028) Ke Bao 2025-08-14 01:41:50 +08:00
  • 4a16a71c36 [PD] feat: mooncake use batch reg/dereg (#8910) Teng Ma 2025-08-14 00:54:34 +08:00
  • a16923efab [PD] optimize kv cache transfer directly using batch transfer (#9149) Francis 2025-08-14 00:54:14 +08:00
  • 6337d9057c [router] optimize Rust compilation and development workflow (#9133) Simo Lin 2025-08-13 05:14:25 -07:00
  • 71fb8c9527 feat: update fa3 (#9126) Yineng Zhang 2025-08-13 05:07:08 -07:00
  • 94f44b88d1 Update fa3 interface and add unit test (#9150) Ke Bao 2025-08-13 20:05:02 +08:00
  • 3b3b3baf9f Double vision prefill throughput by defaulting to optimal vision attention backend (#8484) Kevin Xiang Li 2025-08-13 02:08:30 -07:00
  • 35e6bc92e3 Update docker file for MI35x base image update to support gpt-oss mxfp4 model (#9111) kk 2025-08-13 15:55:31 +08:00
  • 9394ed6386 Fix gpt-oss ~2x memory consumption issue (#9146) fzyzcjy 2025-08-13 15:11:43 +08:00
  • 930fe467bd Support Triton FP8 Gemm can handle hidden_dim not divisible by 16 (#9093) Stefan He 2025-08-12 21:21:55 -07:00
  • 13c48dcf88 [1/2][resubmit again] sgl-kernel: Fuse routed scaling factor into moe_fused_gate (#9088) Trevor Morris 2025-08-12 20:12:38 -07:00
  • 8723b4f146 Use FlashInfer's TRTLLM FP8 Blockscale GEMM (#8588) Elfie Guo 2025-08-12 20:08:40 -07:00
  • 62f99e08b3 fix: wrong docker hub org name (#9137) li chaoran 2025-08-13 10:26:19 +08:00
  • 86a0be65d8 [Feature] Support custom set kv buffer kernel (#8884) DarkSharpness 2025-08-12 16:56:51 -07:00
  • 0edda32001 Support page first layout zero copy for mooncake store (#8651) huangtingwei 2025-08-13 06:59:26 +08:00
  • 924827c3de chore: use cp310 (#9130) Yineng Zhang 2025-08-12 15:33:22 -07:00
  • c81daf838d fix: update Dockerfile (#9129) Yineng Zhang 2025-08-12 15:01:29 -07:00
  • 25caa7a8a9 [AMD] Support Wave attention backend with AMD GPU optimizations (#8660) jacky.cheng 2025-08-13 04:49:11 +08:00
  • 03d114496f Fix typos in supported models documentation (#9119) Hangzhi 2025-08-12 13:35:24 -07:00
  • 83123f481e [Quantization] Supported w8a8 int8 quantized Gemma3 and Qwen-VL models (#8619) ichernob 2025-08-12 23:31:18 +03:00
  • 48afa8f14f [feat] Enable Ascend profiling on SGLang (#8610) ronnie_zheng 2025-08-13 04:28:31 +08:00
  • 2ecbd8b8bf [feat] add ascend readme and docker release (#8700) li chaoran 2025-08-13 04:25:42 +08:00
  • 305b27c124 fix: update Dockerfile (#9125) Yineng Zhang 2025-08-12 13:23:10 -07:00
  • 1ce30dd13e [router] update router documentation (#9121) Simo Lin 2025-08-12 13:16:34 -07:00
  • c9ee738515 Fuse writing KV buffer into rope kernel (part 2: srt) (#9014) Jiaqi Gu 2025-08-12 13:15:30 -07:00
  • 1f9ec65374 fix(docker): update sgl_kernel version to 0.3.4 in Dockerfile.gb200 (#9118) ishandhanani 2025-08-12 13:12:33 -07:00
  • ad359d1c71 router: Fix user guide link README.md (#9122) Chang Su 2025-08-12 12:29:10 -07:00
  • 5f5b3b2449 [5/n] DP Enhancement: Correct num_token_non_padded (#9107) Cheng Wan 2025-08-12 12:23:46 -07:00
  • 4caca4f6b4 Fix typo in REVIEWERS (#9113) Shangming Cai 2025-08-13 02:55:49 +08:00
  • f2a5de284b [Bugfix] Fix accuracy-test-1-gpu failure caused by builtin_tools (#9114) Chang Su 2025-08-12 09:56:13 -07:00
  • 445f9dca6e Runtime check CUDA driver version to avoid unresolved green context symbols (#9021) Liangsheng Yin 2025-08-12 09:26:10 -07:00
  • 3a9afe2a42 chore: bump sgl-kernel v0.3.4 (#9103) Yineng Zhang 2025-08-12 01:48:47 -07:00
  • 9aea255522 Fuse writing KV buffer into rope kernel (part 1: sgl-kernel) (#9077) fzyzcjy 2025-08-12 16:46:40 +08:00
  • fcc11e5ed5 update support new models doc (#9096) Yichao Cheng 2025-08-12 01:21:02 -07:00
  • 5190ba7f42 Fuse two kernels of hidden states padding into quantization kernel (#9005) fzyzcjy 2025-08-12 16:20:13 +08:00
  • 5438886c87 docs: fix broken links in README.md (#9075) Hsiang-Yu Tsou 2025-08-12 15:03:35 +08:00
  • 9c83d74da3 bugfix: Fix the commentary msg extraction in GptOssDetector (#9097) Chang Su 2025-08-11 23:53:10 -07:00
  • b4ac2b9c0c [Fix] Fix dual chunk model default behavior (#9032) DarkSharpness 2025-08-11 23:50:23 -07:00
  • 83262dcb29 Fix mismatch between padded_scales shape and reshape dimensions in modelopt quantization (#8766) Jianwei Dong 2025-08-12 14:44:40 +08:00
  • c46c75f8c0 feat: add fused moe config for Qwen3-30B-A3B on B200 (#9087) zixuanzhang226 2025-08-11 23:25:36 -07:00
  • 2aaf22c46c Optimization for AscendPagedTokenToKVPoolAllocator (#8293) Makcum888e 2025-08-12 09:06:39 +03:00
  • 29a610b4d9 Fix broken CI TestRequestLengthValidation (#9095) Lifu Huang 2025-08-11 22:59:56 -07:00
  • 5ded39cab2 Fix race condition in async lora unload (#9084) Lifu Huang 2025-08-11 22:59:29 -07:00
  • 4093d460ce [CI] migrate router to BM.A10.4 runner (#8992) Keyang Ru 2025-08-11 22:41:18 -07:00
  • 9d68bdb240 [router] Add Rust Binary Entrypoint for SGLang Router (#9089) Simo Lin 2025-08-11 21:37:36 -07:00
  • a218490136 (gpt-oss, oai, chat): Remove Harmony Integration and Implement Native GPT-OSS Tool Call Support (#9043) Chang Su 2025-08-11 18:59:18 -07:00
  • 0eec4cb6cc HiCache, add bench long context plus minor fixs (#9086) Zhiqiang Xie 2025-08-11 16:54:52 -07:00
  • ff1f68252c [fix] Set Radix tree root node hash to None - Nvidia Dynamo Integration (#9030) Faradawn Yang 2025-08-11 14:20:39 -07:00
  • 9f78f391ae HiCache Storage: generate hash when inserting new nodes (#9053) Zhiqiang Xie 2025-08-11 14:18:59 -07:00
  • f508cd3cb7 TRTLLM-MLA FP8 path (#8638) Faraz 2025-08-11 17:02:13 -04:00
  • 44e86480e8 fuse allreduce and residual_rmsnorm (#8731) Xiaoyu Zhang 2025-08-12 04:50:53 +08:00
  • 8c07fabda7 Update hyperparameter_tuning.md (#9083) Lianmin Zheng 2025-08-11 13:44:11 -07:00
  • 90f44b74e6 fix: w4afp8 accuracy problem and rebase (#8752) SijiaYang 2025-08-12 04:41:19 +08:00
  • 38907fe639 refactor(pd-router): extract common patterns to reduce code duplication (#9081) Simo Lin 2025-08-11 13:32:31 -07:00
  • f9afa7dceb Fix docs for clip max new tokens (#9082) Liangsheng Yin 2025-08-11 13:15:21 -07:00
  • 0d9e89ec69 [PD]decode: add CLIP_MAX_NEW_TOKEN for pop_preallocated (#8866) Jimmy 2025-08-12 04:08:11 +08:00
  • 3d64fda376 Fix broken Kimi models HuggingFace link (#9080) Hangzhi 2025-08-11 12:15:00 -07:00
  • 3bffe11279 Fix chunked prefill size validation for disabled state (#8973) 633WHU 2025-08-12 02:05:29 +08:00
  • 44426e54be Update REVIEWERS (#9063) HAI 2025-08-11 11:04:39 -07:00
  • 9f24dfefd1 chore(gb200): remove ToT flashinfer installation (#9079) ishandhanani 2025-08-11 11:02:15 -07:00
  • 89f1d4f536 update deepep commit to support qwen3-coder (#9066) Yi Zhang 2025-08-12 01:42:33 +08:00
  • 75e6a7cde1 Support radix cache for Lora feature (#7216) Baizhou Zhang 2025-08-11 10:14:11 -07:00
  • 6f81a710f7 [pd-router] add retry and circuit breakfor for pd router (#9051) Simo Lin 2025-08-11 05:53:26 -07:00
  • a6452b7188 bugfix: Fix output_ids extraction in detokenizer_manager (#9047) Chang Su 2025-08-11 03:17:32 -07:00
  • f4ae50e97c fix: use flashinfer v0.2.11.post1 zhyncs 2025-08-11 02:49:25 -07:00
  • 84cb449eec Revert "chore: upgrade flashinfer 0.2.11 (#9036)" (#9057) Yineng Zhang 2025-08-11 00:16:39 -07:00
  • f003cd3548 [CI] Fix CI tests (#9050) Cheng Wan 2025-08-10 23:52:05 -07:00
  • 9d834fdcc1 Revert "feat: update flashinfer ar oneshot params (#8687)" (#9054) Yineng Zhang 2025-08-10 23:24:42 -07:00
  • b32792516a REVIEWERS.md typo fix (#9048) Zhiqiang Xie 2025-08-10 22:33:37 -07:00
  • 067068f271 [router] regular router circuit breaker (#8997) Simo Lin 2025-08-10 21:19:30 -07:00
  • 6beeff41c5 Update REVIEWERS.md (#9046) Lianmin Zheng 2025-08-10 21:11:14 -07:00
  • 2e8e7e353b Improve docs and developer guide (#9044) Lianmin Zheng 2025-08-10 21:05:18 -07:00
  • 2449a0afe2 Refactor the docs (#9031) Lianmin Zheng 2025-08-10 19:49:45 -07:00
  • 0f229c07f1 Update release-docs.yml (#9037) Lianmin Zheng 2025-08-10 18:52:11 -07:00
  • dd001a5477 chore: upgrade flashinfer 0.2.11 (#9036) Yineng Zhang 2025-08-10 17:35:37 -07:00