Commit Graph

  • a69b637014 [router] fix req handling order, improve serialization, remove retry (#8888) Simo Lin 2025-08-06 23:24:39 -07:00
  • 2d120f8b18 [Feature][Multimodal] Implement LRU cache for multimodal embeddings (#8292) Zheng Wengang 2025-08-07 14:21:40 +08:00
  • 4f2e1490c3 [AMD] Pull latest SGLang version for AMD CI (#8787) michael-amd 2025-08-06 20:20:26 -07:00
  • 3fa3c6cd6a Enables force reasoning based on chat template for Qwen3-Thinking (#8369) Xinyuan Tong 2025-08-06 20:02:47 -07:00
  • 6210e2c4f0 Support GPU pinning for LoRA (#8697) Lifu Huang 2025-08-06 19:39:45 -07:00
  • 6ad6c8c9e6 feat: openai oss attention sink support with trtllm-gen backend #8825 (#8834) eigen 2025-08-06 22:18:27 -04:00
  • 5b6acc1495 fix glm4 moe (#8883) Cheng Wan 2025-08-06 18:02:31 -07:00
  • 4373df5525 add flashinfer mxfp4 (#8847) Xiaoyu Zhang 2025-08-07 07:23:41 +08:00
  • c0e84297c2 Use reduce scatter for DP (#8539) Trevor Morris 2025-08-06 16:21:26 -07:00
  • 92cc32d9fc Support v1/responses and use harmony in serving_chat (#8837) Chang Su 2025-08-06 16:20:34 -07:00
  • cbbd685a46 chore: use torch 2.8 stable (#8880) Yineng Zhang 2025-08-06 15:51:40 -07:00
  • 78aad91037 [CI] fix pip upgrade (#8881) Cheng Wan 2025-08-06 15:02:32 -07:00
  • 288ae41f7a [NVIDIA] Fix num_experts in modelopt_quant (#8811) Shu Wang 2025-08-06 16:35:07 -05:00
  • 01c99a9959 chore: update Dockerfile (#8872) Mick 2025-08-07 00:30:33 +08:00
  • b114a8105b Support B200 in CI (#8861) fzyzcjy 2025-08-06 21:42:44 +08:00
  • 0475448ee3 Optimize triton swa kernel by skipping computation (#8860) Ke Bao 2025-08-06 21:37:50 +08:00
  • 399e7ec8b3 Refine naming (#8868) Ke Bao 2025-08-06 21:37:02 +08:00
  • 1bd5316873 fix benchmark fp8 blockwise group gemm (#8815) Yuan Luo 2025-08-06 21:02:21 +08:00
  • aeac900ca2 fix: resolve ci issue (#8859) Yineng Zhang 2025-08-06 02:28:14 -07:00
  • 4fc5f2f977 Add unit test for triton swa kernel (#8853) Ke Bao 2025-08-06 16:10:38 +08:00
  • 168033d5fb Support mxfp4 for GPT-OSS (#8843) Ying Sheng 2025-08-06 00:05:25 -07:00
  • cbbb738371 [2/3] Optimize Slime Update Weights: Avoid GPU-to-CPU Device Sync when update expert weights (#8753) Stefan He 2025-08-05 22:09:52 -07:00
  • 89588179cf [1/3] Optimize Slime Update Weights: Remove QWen3MOE Load Weight Overhead (#8751) Stefan He 2025-08-05 22:07:54 -07:00
  • 8c7bb39dfb [router] PD Router Simplification and Reorganization (#8838) Simo Lin 2025-08-05 21:20:38 -07:00
  • ca47e24f5d [Feature] improve TBO: two chunk overlap (#8144) HouseWest 2025-08-06 12:11:01 +08:00
  • d26ca84f39 Support bailing moe (#8680) Praneth Paruchuri 2025-08-06 09:10:34 +05:30
  • 8128e08d36 Turn off hybrid cache by default (#8839) Ke Bao 2025-08-06 09:53:45 +08:00
  • 5d62b56f7e [router] complete router oai spec (#8828) Simo Lin 2025-08-05 18:30:19 -07:00
  • 3ae8e3ea8f chore: upgrade torch 2.8.0 (#8836) Yineng Zhang 2025-08-05 17:32:01 -07:00
  • c1d2061f97 Add initial support for gpt-oss (#8824) Ying Sheng 2025-08-05 13:42:01 -07:00
  • 556e4143f0 fix: remove unused import (#8809) Yineng Zhang 2025-08-05 13:40:22 -07:00
  • 4ef47839ae feat: use py312 (#8832) Yineng Zhang 2025-08-05 13:38:22 -07:00
  • 32d9e39a29 Fix potential memory fault issue and ncclSystemError in CI test (#8681) kk 2025-08-06 03:19:37 +08:00
  • 4f4e0e4162 chore: upgrade flashinfer 0.2.10 (#8827) Yineng Zhang 2025-08-05 12:04:01 -07:00
  • 901ab758ec chore: upgrade transformers 4.55.0 (#8823) Yineng Zhang 2025-08-05 11:37:21 -07:00
  • 8e8545caf6 fix: update cmake (#8817) Yineng Zhang 2025-08-05 09:38:30 -07:00
  • a4b0d5c9e5 GLM-4.5 and GLM-4.5-Air both support (#8804) Yuxuan Zhang 2025-08-05 18:29:20 +08:00
  • 40e3b2beeb feat: add trtllm-gen mha from direct call (#8782) eigen 2025-08-05 06:28:39 -04:00
  • 75df31b60e chore: bump sgl-kernel v0.3.2 (#8802) Yineng Zhang 2025-08-05 02:35:20 -07:00
  • 194561f27a feat: support sgl-kernel cu129 (#8800) Yineng Zhang 2025-08-05 02:33:47 -07:00
  • 5e91fed1c5 Revert "[NVIDIA]Fix local_num_experts for EP (#8779)" (#8797) Yineng Zhang 2025-08-04 23:30:43 -07:00
  • 873f384a51 [feat] Add detail in image_data (#8596) Yuhao Yao 2025-08-05 14:01:38 +08:00
  • b01eeb80f8 [NVIDIA]Fix local_num_experts for EP (#8779) Shu Wang 2025-08-05 00:01:14 -05:00
  • 1ea94d3b92 chore: upgrade flashinfer v0.2.9 (#8780) Yineng Zhang 2025-08-04 21:59:18 -07:00
  • 354ac43555 [pd-router] Add Configurable Retry Logic for reduce backend pressure (#8744) Simo Lin 2025-08-04 20:42:07 -07:00
  • d98a4913ea [PD] Refactor parallel sizes and add pp support for mooncake (#8571) Shangming Cai 2025-08-05 11:18:11 +08:00
  • 08f8f49016 [CPU][sgl-kernel] biased_grouped_topk: fix correction_bias dtype to float32 (#8212) Chunyuan WU 2025-08-05 09:28:31 +08:00
  • d4bf5a8524 Support OCP MXFP4 quantization on AMD GPUs (#8255) kk 2025-08-05 09:14:52 +08:00
  • 7cb20754fa [Fix] Fix several issues preventing gemma3n LoRA support. (#8776) Lifu Huang 2025-08-04 17:11:46 -07:00
  • 6d0646da11 [NVIDIA] Fix breakage of using trtllm-gen fp8 moe (#8773) Kaixi Hou 2025-08-04 16:30:13 -07:00
  • 02bc1c7d80 chore: bump sgl-kernel v0.3.1 (#8771) Yineng Zhang 2025-08-04 13:18:54 -07:00
  • fc8c8e5041 Integrate triton_kernels in sgl-kernel (#8762) Qiaolin Yu 2025-08-04 12:12:14 -07:00
  • 9bd4872a34 [bugfix] Fix typo in modelopt quant: 'FusedMoE' object has no attribute 'local_num_experts' (#8768) Trevor Morris 2025-08-04 11:08:08 -07:00
  • 2fa0462c39 [router] introduce dp worker abstraction (#8639) Simo Lin 2025-08-04 06:42:20 -07:00
  • 915140fd18 [NVIDIA] Add Low Latency NVFP4 decode kernels from Flashinfer (#8552) azhurkevich 2025-08-04 03:10:02 -07:00
  • 36fc9260a2 [bugfix] fix import path in HiCacheController (#8749) Baron Liu 2025-08-04 13:19:15 +08:00
  • fee0ab0fba [CI] Ascend NPU CI enhancement (#8294) Even Zhou 2025-08-04 13:16:38 +08:00
  • f57d2dc162 [sgl-kernel] avoid per_token_quant_fp8.cu hardcode sm_count (#8738) Xiaoyu Zhang 2025-08-04 12:55:57 +08:00
  • f2d68ded6d Rename lora_path to lora_id in batches (#8437) Baizhou Zhang 2025-08-03 21:08:28 -07:00
  • 3b87a9e8ae Fix bug of refactoring TopKOutput in w4afp8 (#8745) Yuan Luo 2025-08-04 11:05:02 +08:00
  • f024795e57 Replace torch.jit.script with torch.compile in get_masked_input_and_mask to fix benchmark underreporting (#8733) YyWangCS 2025-08-04 10:02:51 +08:00
  • b102353f8f [MoE] Enable renormalize=False in Triton kernels (#8735) Cheng Wan 2025-08-03 17:03:04 -07:00
  • 7a27e798ca [CI] Do not trigger pd-disaggregation CI in draft PR (#8737) Liangsheng Yin 2025-08-04 05:12:20 +08:00
  • 76ba5bbe12 fix args typo in memory_pool_host (#8662) huangtingwei 2025-08-04 04:47:29 +08:00
  • ed6f7597b3 Fix the missing 'lof' choice of --schedule-policy server args (#7114) Yingchun Lai 2025-08-04 03:29:42 +08:00
  • e67276ecb3 feat: support cutlass_moe_fp8 kernel for fusedmoe in sm90 (#8678) tql.99 2025-08-04 01:47:15 +08:00
  • 0242bb9c74 Fix triton kernels topk with keyword arguments (#8732) Ke Bao 2025-08-04 01:45:15 +08:00
  • 760286e3d3 use fp32 for e_score_correction_bias in GLM-4.5 (#8729) Yuxuan Zhang 2025-08-04 01:43:40 +08:00
  • 3435a24e81 [RL] fix update weight for FusedMoE with EP (#8676) Zilin Zhu 2025-08-04 01:20:39 +08:00
  • 00da906584 feat: Support DP Attention for step3_vl (#8699) yhyang201 2025-08-03 19:35:26 +08:00
  • 8cd344586e chore: bump v0.4.10.post2 (#8727) Yineng Zhang 2025-08-03 03:43:29 -07:00
  • 0e0eef00ce [DP] fix the compatibility issue between DP attention and --attention-backend triton (#8723) Cheng Wan 2025-08-03 03:06:57 -07:00
  • cb099d2095 [CUDA Graph] save cuda graph memory by using next_token_logits_buffer (#8579) Cheng Wan 2025-08-03 03:06:47 -07:00
  • 7a91330149 Save cuda graph memory for fa3 (#8567) Cheng Wan 2025-08-03 03:06:31 -07:00
  • 5ce5093b97 chore: bump sgl-kernel 0.3.0 with torch 2.8.0 (#8718) Yineng Zhang 2025-08-03 02:31:50 -07:00
  • 6f9baf1002 [Improvements] Merge health check route (#8444) ybyang 2025-08-03 16:59:06 +08:00
  • a31b7a7024 feat: Add new moe triton for NVIDIA RTX 6000 Ada (#8547) Jasper James 2025-08-03 15:57:35 +08:00
  • 7ed8e51bc3 [fix] Fix divide by zero error for llama4. (#8683) Varun Vinayak Shenoy 2025-08-03 00:55:55 -07:00
  • 32f2815451 Do layernorm before allgather for DP attention (#8631) Trevor Morris 2025-08-03 00:53:08 -07:00
  • f7b2853ff8 [feat] support minimum token load balance in dp attention (#7379) Guanhua Wang 2025-08-03 15:46:47 +08:00
  • b0add2da00 HiCache storage, style change and bug fix (#8719) Zhiqiang Xie 2025-08-03 00:05:04 -07:00
  • 0305c5053f Reduce memory accumulation in long-running server (#8306) Wenxuan Tan 2025-08-03 02:03:16 -05:00
  • 8675bdf246 Support limiting max loaded loras in CPU. (#8650) Lifu Huang 2025-08-03 00:02:23 -07:00
  • a437aa9987 [hotfix] fix mixtral with tensor-level compressed-tensor quantization (#8721) Cheng Wan 2025-08-02 22:59:25 -07:00
  • 0e612dbf12 Tiny fix CI pytest error (#8524) fzyzcjy 2025-08-03 13:48:42 +08:00
  • 9f47d686e5 Fix fused MoE when routed_scaling_factor is None (#8709) Liangsheng Yin 2025-08-03 12:42:01 +08:00
  • d9def43dcd [Perf]Use Cooperative Schedule for H100 & H200 & H800 in fp8_blockwise_scaled_grouped_mm (#8722) Qi Yuhang 2025-08-03 12:13:47 +08:00
  • e273aa6dcf [Feature] Radix Tree in C++ (#7369) DarkSharpness 2025-08-02 19:50:14 -07:00
  • 828a4fe944 [router] Implement HTTP Dependency Injection Pattern for Router System (#8714) Simo Lin 2025-08-02 19:16:47 -07:00
  • 8ada1ab6c7 Fix triton moe error caused by TopK refactor (#8705) fzyzcjy 2025-08-03 09:49:47 +08:00
  • e314b084c5 [FIX] Fix the nightly CI by disabling swa mem pool for gemma2 (#8693) Lianmin Zheng 2025-08-02 18:43:14 -07:00
  • 403566bcca Remove assertions about per group quant fp8 (#8717) fzyzcjy 2025-08-03 08:08:40 +08:00
  • 0a56b721d5 chore: bump sgl-kernel v0.2.9 (#8713) Yineng Zhang 2025-08-02 16:21:56 -07:00
  • 603f5ce020 [Bug] fix green context's incompatibility with cuda < 12.4 (#8701) Liangsheng Yin 2025-08-03 06:23:11 +08:00
  • 6d4fd8826e [router] minor code clean up and and refactoring (#8711) Simo Lin 2025-08-02 13:46:31 -07:00
  • f9f0138f80 Revert "[1/2] sgl-kernel: Fuse routed scaling factor into select_experts" (#8706) Liangsheng Yin 2025-08-02 20:14:30 +08:00
  • ac6962ccd6 [Doc] Polish sgl-kernel readme for cu126 build error (#8704) PGFLMG 2025-08-02 17:03:07 +08:00
  • 4ca43b061c Add tensor.detach() back to update weight util (#8691) Stefan He 2025-08-02 00:41:05 -07:00
  • ea93079b30 model: adapt mllama4 to VisionAttention (#8512) Wenchen Lo 2025-08-02 00:39:40 -07:00
  • 4bec99ecd0 Fix: resolve prefill of retracted request out-of-memory issue when ignore_eos is enabled (#7434) Yusong Gao 2025-08-02 14:43:45 +08:00