Commit Graph

  • 4ed67c27e3 [router] support Openai router conversation API CRUD (#11297) Keyang Ru 2025-10-07 15:31:35 -07:00
  • cd4b39a900 [quantization] Properly ignore quantization for layers excluded in quant_config (#11205) Bowen Bao 2025-10-07 14:06:05 -07:00
  • 420c99acfe [router][grpc] Fix error message format in grpc chat handler (#11307) Chang Su 2025-10-07 13:54:02 -07:00
  • e3c7f09146 Update tool parser and related documentation (#11223) Xinyuan Tong 2025-10-07 11:03:40 -07:00
  • 6f1e03a456 [router][grpc] Fix sampling_params.stop_strs is None (#11306) Chang Su 2025-10-07 10:57:38 -07:00
  • f4affd4df5 [router] fix grpc connection conversion and add optimization (#11305) Simo Lin 2025-10-07 13:39:33 -04:00
  • df08bf9b9f [Doc]: Best Practice for HICache (#11001) hzh0425 2025-10-08 00:59:21 +08:00
  • 69efdd27bc [Doc] HiCache Design Documents (#11027) ykwd 2025-10-08 00:35:45 +08:00
  • 64582caa84 [router][grpc] Refactor chat template content format detection (#11288) Chang Su 2025-10-07 08:38:51 -07:00
  • 2fcd56eaf6 [router] add get server info and get model info in grpc server (#11303) Simo Lin 2025-10-07 11:36:52 -04:00
  • 0958a39704 [Docs] [Router] Update Observability and Common Issues Section (#11302) Wenyi Xu 2025-10-07 23:03:09 +08:00
  • 4f42c8cd3e [sgl-kernel] Support float64 moe_sum_reduce cuda kernel (#11068) Yuan Luo 2025-10-07 22:31:11 +08:00
  • 3ddd7dc9f8 Introduce future indices (#11301) Liangsheng Yin 2025-10-07 22:24:02 +08:00
  • 501dfa6b42 Remove sampling info events and overlap thread file (#11300) Liangsheng Yin 2025-10-07 21:34:25 +08:00
  • 79d3495177 [router] add reasoning and tool parser argument in router (#11290) Simo Lin 2025-10-07 09:08:32 -04:00
  • 1519a89cfd Remove overlap thread (#11210) Liangsheng Yin 2025-10-07 20:12:12 +08:00
  • 24bc3fb0f9 EAGLE cache fix for SWARadixCache (#11231) Ke Bao 2025-10-07 18:21:37 +08:00
  • 8a8a608af9 [ci] fix pp test (#11294) Liangsheng Yin 2025-10-07 14:20:04 +08:00
  • 533e58a15d Feature/longbench v2 evaluation utils (#10949) Al-Ekram Elahee Hridoy 2025-10-07 00:17:31 -06:00
  • 9b4c449735 convert test_deterministic into unit tests (#11095) Alex Chi Z 2025-10-07 05:33:11 +02:00
  • a578d300ba [router][grpc] Fix proto3 default value mismatches and cleanup unused fields (#11283) Chang Su 2025-10-06 18:54:51 -07:00
  • 8c9670375f chore: bump sgl-kernel version to 0.3.15 (#11281) sglang-bot 2025-10-06 18:17:51 -07:00
  • fb27d38305 docs: update sgl-kernel README (#11286) Yineng Zhang 2025-10-06 17:55:22 -07:00
  • fd8a0b29c0 fix: correct scale parameter remapping logic in Llama4ForConditionalGeneration (#11282) Xinyuan Tong 2025-10-06 17:28:23 -07:00
  • afc35ccc5e Fix LoRA support for multimodal models (VLMs) by implementing a consistent pattern for skipping vision components (#11261) Chenxi Li 2025-10-06 17:23:00 -07:00
  • a57f0e3d56 reverse the amd ci test back to 1200s and split the 8-gpu deepseek job into two. (#11238) sunxxuns 2025-10-06 19:27:57 -04:00
  • 708f4ff490 Rename max_micro_batch_size -> pp_max_micro_batch_size (#11279) Lianmin Zheng 2025-10-06 15:50:56 -07:00
  • e2daeb351c [Auto Sync] Update test_utils.py (20251006) (#11280) Lianmin Zheng 2025-10-06 15:49:57 -07:00
  • 0e7b353009 Fix code sync scripts (#11276) Lianmin Zheng 2025-10-06 15:35:01 -07:00
  • b07c9c76c5 [router][grpc] Refine streaming processes (#11277) Chang Su 2025-10-06 15:15:01 -07:00
  • 748f86f3de [Bug] Fix incorrect assertion in FA4 and add UT. (#11182) Lifu Huang 2025-10-06 14:58:39 -07:00
  • 73ea484af1 docker: add manifest to versioned docker releases (#11268) ishandhanani 2025-10-06 14:53:40 -07:00
  • 4aeb193fbd disable sm100 for FlashMLA and fast-hadamard-transform in cuda12.6.1 (#11274) gongwei-130 2025-10-06 14:48:31 -07:00
  • 466992b2d0 [router][tool call] Clean up redundant detect_format and has_tool_markers (#11270) Chang Su 2025-10-06 14:04:02 -07:00
  • 155cbb51f0 Enable native ModelOpt quantization support (1/3) (#7149) Zhiyu 2025-10-06 13:24:15 -07:00
  • eb30b888db Remove env var warnings for release (#11262) Lianmin Zheng 2025-10-06 10:09:17 -07:00
  • 5ee777c98f [router] add ipv6 support across all components (#11219) Simo Lin 2025-10-06 11:16:59 -04:00
  • a4a3d82393 chore: bump SGLang version to 0.5.3 (#11263) sglang-bot 2025-10-06 05:07:02 -07:00
  • 0b13cbb7c9 chore: bump SGLang version to 0.5.3rc2 (#11259) sglang-bot 2025-10-06 01:12:10 -07:00
  • efbc687c28 Support DeepSeek V3.2 Exp (#11061) fzyzcjy 2025-10-06 15:24:15 +08:00
  • 292a867ad9 Add flashmla and fast hadamard transform to Dockerfile (#11235) Baizhou Zhang 2025-10-05 21:31:28 -07:00
  • 8fd41eae93 Improve bot release workflow (#11240) Kangyan-Zhou 2025-10-05 21:28:27 -07:00
  • 0cd1996eae feat: add shortcut detection for multimodal templates in Jinja format (#11209) Xinyuan Tong 2025-10-05 21:13:17 -07:00
  • b6b4b56395 Update condition for sgl-kernel-benchmark-test (#11254) Lianmin Zheng 2025-10-05 20:55:02 -07:00
  • f8924ad74b update sgl kernel version to 0.3.14.post1 (#11242) Lianmin Zheng 2025-10-05 20:30:40 -07:00
  • 2f80bd9f0e Bump torch_memory_saver 0.0.9rc2 (#11252) fzyzcjy 2025-10-06 11:26:20 +08:00
  • 366a603e95 Use cu128 for torch audio to fix some CI tests (#11251) Lianmin Zheng 2025-10-05 19:52:32 -07:00
  • baee08601b [quantization] Enable aiter mxfp4 fused_moe for Quark (#10048) Bowen Bao 2025-10-05 19:51:34 -07:00
  • c7a104c12b [quantization] Fix scale remapping for mllama4 (#10042) Bowen Bao 2025-10-05 19:51:15 -07:00
  • 97d966a7f8 ci: make find_local_hf_snapshot_dir more robust (#11248) Mick 2025-10-06 10:50:11 +08:00
  • 8e66d87f0a Fix spec_utils.py (#11247) sglang-bot 2025-10-05 19:01:11 -07:00
  • a20fc7b7dc Create two new GH workflows to automatically bump SGLang and Kernel version (#10996) Kangyan-Zhou 2025-10-05 18:14:05 -07:00
  • 6b30e097ab [Auto Sync] Update io_struct.py (20251004) (#11206) Lianmin Zheng 2025-10-05 18:06:07 -07:00
  • d645ae90a3 Rename runner labels (#11228) Lianmin Zheng 2025-10-05 18:05:41 -07:00
  • 41763ba079 Remove gdrcopy check in ci_install_deepep.sh (#11237) Cheng Wan 2025-10-05 17:35:22 -07:00
  • 652c24a653 Update transformers package version to 4.57.0 (#11222) Xinyuan Tong 2025-10-05 16:45:14 -07:00
  • 5e142484e2 [Fix AMD CI] VRAM cleanup (#11174) sunxxuns 2025-10-05 19:03:53 -04:00
  • c560410da7 Refactor and optimize mooncake CI (#11162) Shangming Cai 2025-10-06 05:08:52 +08:00
  • 590f2da052 [Feat] Support Torch Symm Mem AllReduce (#10571) Yuan Luo 2025-10-06 04:55:19 +08:00
  • 148d8d485d Update DeepGEMM repository tag to specific commit (#11229) Lianmin Zheng 2025-10-05 13:47:36 -07:00
  • 1a599509cc chore: bump sgl-kernel v0.3.14.post1 (#11137) PGFLMG 2025-10-06 04:46:43 +08:00
  • 36a6b8dbfc Update v1/responses to be more OpenAI-compatible. (#9624) Vincent Zhong 2025-10-05 14:47:46 -04:00
  • e0b2d3eebe [Feature] Add a fast-topk to sgl-kernel for DeepSeek v3.2 (#11194) DarkSharpness 2025-10-06 01:19:03 +08:00
  • 4cb5a5235e Tiny skip_sample adjust (#11225) Liangsheng Yin 2025-10-05 23:41:04 +08:00
  • 85c1f79377 Add DeepSeek-V3.2 Tool Call Template (#11063) Xu Wenqing 2025-10-05 09:53:49 +08:00
  • 48e9e71930 Add --max-new-tokens CLI flag for MMMU evaluation (#11217) yhyang201 2025-10-05 08:35:53 +08:00
  • 31b49c0b51 EAGLE cache fix for HiCache (#11215) Ke Bao 2025-10-05 07:53:53 +08:00
  • d736e0b65e [router] add grpc router pd mode for chat and generate (#11140) Simo Lin 2025-10-04 09:58:28 -04:00
  • ffd03a9bd3 [router] fix get load response parsing (#11213) Simo Lin 2025-10-04 09:58:02 -04:00
  • 666da3d59f [fix]enable flashmla when using draft model P/D attention select (#11012) Hank Han 2025-10-04 20:59:34 +08:00
  • d01b921482 fix sampling_seed handling when deterministic is enabled (#11096) Alex Chi Z 2025-10-03 23:41:46 -04:00
  • c70e58e837 [HICache]: Refactor HiCache CI (#11011) hzh0425 2025-10-04 08:51:56 +08:00
  • c61b9a1d01 fix self.enable_kv_cache_events (#11178) narutolhy 2025-10-03 14:09:41 -07:00
  • 3c3d6255d9 [fix]missing prefix_lens_cpu init when p/d disaggregation (#11196) Hank Han 2025-10-04 04:39:59 +08:00
  • 546914fa2d [Fix] Fix the bug of the calculation of base_gpu_id (dp offset) in data_parallel_controller.py (#10741) XSongQ 2025-10-03 13:25:57 -07:00
  • 4726c9197f [minor] fix the lint (#11198) Liangsheng Yin 2025-10-04 01:04:58 +08:00
  • a0010bf4e8 fix qwen2 eagle3 runtime error (#10517) jiapingW 2025-10-04 00:19:52 +08:00
  • 307fc060e8 fix xeon ci check (#10838) DiweiSun 2025-10-04 00:17:36 +08:00
  • 586e81a28a [Test] Initialize mem_fraction_static in setUpClass to fix pytest VLM test crashes. (#10859) vikram singh shekhawat 2025-10-03 21:44:48 +05:30
  • fad7ca73f8 model: support starcoder2 (#10609) Praneth Paruchuri 2025-10-03 21:41:19 +05:30
  • 08af8ffb5c fix 3fs indices (#10855) pansicheng 2025-10-04 00:06:38 +08:00
  • 2c7f4ca2f2 Optimize debug log position of PD abort request (#11090) Shangming Cai 2025-10-03 23:07:02 +08:00
  • 03def5e3b1 Fix [test]: Env:SGLANG_TORCH_PROFILER_DIR for pytest. (#10780) shubham singhal 2025-10-03 20:29:32 +05:30
  • 6ae3f05b33 Fix CUDA illegal memory access issues in speculative decoding (#10892) ur4t 2025-10-03 22:44:07 +08:00
  • fdc4e1e570 Tiny move files to utils folder (#11166) fzyzcjy 2025-10-03 22:40:06 +08:00
  • 04b86b3c5c [hot-fix] Fix CI break which caused by adding thinking_mode in eval (#11192) Liangsheng Yin 2025-10-03 18:29:27 +08:00
  • d6777a706d Add --thinking-mode to run_eval (#11189) hlu1 2025-10-03 01:49:39 -07:00
  • 8c57490210 [Feature] Option to save model weights to CPU when memory saver mode is enabled (#10873) Matt Nappo 2025-10-03 04:48:19 -04:00
  • 34151f173b [router] Steaming support for MCP Tool Calls in OpenAI Router (#11173) Keyang Ru 2025-10-03 00:19:43 -07:00
  • 6794d21051 Tiny add PD disaggregation + DP attention test (#11167) fzyzcjy 2025-10-03 14:15:46 +08:00
  • 1a31229cd4 fix: radix cache memory accounting (#10637) Alex Chi Z 2025-10-03 01:47:33 -04:00
  • de89ef49da [CI]] Tee server logs to both file and stdout/stderr using PIPE (#11185) Liangsheng Yin 2025-10-03 12:31:13 +08:00
  • b00a0c786f [Fix] Update to v0.1.5.post4 and refine HIP attention backend selection (#11161) jacky.cheng 2025-10-03 12:19:30 +08:00
  • a2faf8940c [1/n] Enable DCA CUDA graph capture (#9537) b8zhong 2025-10-02 20:30:00 -07:00
  • 7e61737d3f [Generative Scores API] add performance tests to CICD (#10830) Vedant V Jhaveri 2025-10-02 19:57:55 -07:00
  • 3c699772c9 Introduce naming convention in io_struct and base sglang io classes. (#10133) Liangsheng Yin 2025-10-03 10:55:13 +08:00
  • e810077488 Allow use of TRTLLM_MHA backend for hybrid attention on Blackwell (#11138) Dom Brown 2025-10-03 00:04:58 +01:00
  • 963175d5c0 [router][grpc] Support streaming for v1/chat/completions (#11179) Chang Su 2025-10-02 14:35:16 -07:00
  • 0618ad6dd5 fix: shoudn't include CUDA_ARCH 100 and 120 for cuda12.6.1 (#11176) gongwei-130 2025-10-02 13:24:23 -07:00
  • 6a261aaca5 Minor fixes for server_args, parallel_state, and test_deterministic.py (#11159) Lianmin Zheng 2025-10-02 12:12:49 -07:00