Commit Graph

  • 839c93bd2d feat: add original logprobs to response (#8375) narutolhy 2025-08-29 11:43:57 -07:00
  • f1e9bbaff5 feat: Add flexible validation for partial weight updates (#9663) JiLi 2025-08-30 02:19:26 +08:00
  • 3fd1431df2 support enable in the reasoning field to enable thingking for thinkin… (#9715) gongwei-130 2025-08-29 10:57:32 -07:00
  • 161e9dc51e feat(hicache-3fs): 3FS-Store Backup Optimizations For MLA Model. (#9692) hzh0425 2025-08-30 01:48:51 +08:00
  • 54e872d343 [HiCache] resolve conflict between chunked-prefill and hicache hit count (#9776) Zhiqiang Xie 2025-08-29 10:30:54 -07:00
  • e5b29bf14e [PD] Support get_model_info interface for mini_lb (#9792) Xuchun Shang 2025-08-29 15:54:03 +08:00
  • 9a7c8842ba accomendate json schema in the "schema" field, not in "json_schema" field of response_format (#9786) gongwei-130 2025-08-28 23:51:50 -07:00
  • 7a16db9bd9 Make sm100 fp8 kernels available on sm103 (#9789) hlu1 2025-08-28 23:47:29 -07:00
  • 09a1df2231 add bench_mix.py (#9788) pansicheng 2025-08-29 14:44:26 +08:00
  • 4b7034ddb0 ROCm 7.0 update (#9757) sogalin 2025-08-28 22:24:34 -07:00
  • a23c30205d Raise error when topk>1 and page>1 for paged attention backends. (#9784) Liangsheng Yin 2025-08-29 12:47:34 +08:00
  • a7d825fccc Skip some tests on Blackwell (#9777) hlu1 2025-08-28 20:00:32 -07:00
  • 38cd5fb1e0 bugfix(hicache): Move exists check before key suffixing (#9749) hzh0425 2025-08-29 09:29:47 +08:00
  • 001f51940a [HiCache] change the default policy to write through (#9772) Zhiqiang Xie 2025-08-28 18:28:39 -07:00
  • 5ad296bda1 Optimize prefill performance on cpu backend (#8750) Ma Mingfei 2025-08-29 08:21:55 +08:00
  • 9f81d741a2 fix: fix MLA for ShardedModelLoader/RemoteModelLoader (#6287) wangyu 2025-08-29 07:10:09 +08:00
  • a38c149758 feat(draft_model): support draft_model for RemoteModelLoader (#6407) wangyu 2025-08-29 07:09:52 +08:00
  • 74dd4249ac [Feature] Support NPUGraph for DeepSeek on Ascend NPU (#9355) chenxu140 2025-08-29 07:06:24 +08:00
  • dc20c22f76 feat: add tuned fused moe config for GLM-4.5-Air-FP8 tp = 4 on B200 (#9770) zixuanzhang226 2025-08-28 16:00:28 -07:00
  • 711390a971 [AMD] Support Hierarchical Caching on AMD GPUs (#8236) Hubert Lu 2025-08-28 15:27:07 -07:00
  • 5343058875 [router] grpc router bootstraps (#9759) Simo Lin 2025-08-28 12:07:06 -07:00
  • fce7ae33f8 [Sync] Update server_args.py (20250828) (#9745) Lianmin Zheng 2025-08-28 10:33:00 -07:00
  • 6b39f9cf8c Support compile sgl-kernel on cuda 13.0 (#9721) Rain Jiang 2025-08-28 10:18:03 -07:00
  • 07c9d8fba2 [router] add llama3.2 multi json streaming parser (#9735) Simo Lin 2025-08-28 05:57:13 -07:00
  • 4a4772ae03 Support speculative decoding in hybrid attention backend (#9573) Qiaolin Yu 2025-08-28 01:11:42 -07:00
  • c377923304 [feat] Reduce GPU memory overhead by using weakref (#9673) yhyang201 2025-08-28 16:09:06 +08:00
  • f84b57c80e Move git clone command up from README (#9740) Xinyuan Tong 2025-08-28 07:27:00 +00:00
  • aee094e430 add support for nvidia/gpt-oss-120b-Eagle3 (#9739) zyksir 2025-08-28 15:20:20 +08:00
  • 55349e361d support mooncake store dp attention (#9684) huangtingwei 2025-08-28 12:31:31 +08:00
  • e1f7cf57dc [router] additional llama32 parser unit test and multi json support (#9732) Simo Lin 2025-08-27 20:34:11 -07:00
  • 2bb9d454b5 [router] additional pythonic parser unit test (#9730) Simo Lin 2025-08-27 19:55:59 -07:00
  • d0934a5192 gpt-oss blog reproduction document (#9728) Liangsheng Yin 2025-08-28 10:15:08 +08:00
  • 3f2d0cefcd [router] Add MCP Tool Handler (#9615) Keyang Ru 2025-08-27 19:12:39 -07:00
  • 8b30bec265 [router] fix error response in pd_router (#9505) Bruce-x-1997 2025-08-28 10:10:55 +08:00
  • 4aeba40d7b [Sync] Update mxfp4.py (20250827) (#9724) Lianmin Zheng 2025-08-27 17:00:09 -07:00
  • 28684f909d [router] upgrade kernel version in pd ci (#9720) Chang Su 2025-08-27 16:02:41 -07:00
  • bc80dc4ce0 chore: bump v0.5.1.post3 (#9716) Yineng Zhang 2025-08-27 15:42:42 -07:00
  • b962a296ed chore: upgrade sgl-kernel 0.3.7 (#9708) Yineng Zhang 2025-08-27 14:00:31 -07:00
  • aa3eba8eb4 [sgl-kernel] misc: update deepgemm version for sgl-kernel (#9340) PGFLMG 2025-08-28 03:01:30 +08:00
  • 07ee0ab750 [router] add gpt-oss and glm4 tool parser (#9703) Simo Lin 2025-08-27 11:26:00 -07:00
  • 5c06dcb75a [router] add kimi-k2 tool parser (#9702) Simo Lin 2025-08-27 11:04:55 -07:00
  • 6f6beca49d [router] add step3 tool parser (#9695) Simo Lin 2025-08-27 10:44:52 -07:00
  • 68a54e063e Sets default model name in request classes (#9683) Xinyuan Tong 2025-08-27 17:43:03 +00:00
  • fd18995cf3 Fix get_ip when no external network (#9700) ybyang 2025-08-28 01:28:52 +08:00
  • db0831e019 Quick fix for loading processor for supporting internvl3_5 series (#9676) yilian49 2025-08-27 12:05:27 -04:00
  • 6e4e1c8cdc [router] add deepseek tool parser (#9694) Simo Lin 2025-08-27 06:18:24 -07:00
  • 9768c50d90 [router] restructure tool parser module folder (#9693) Simo Lin 2025-08-27 06:05:53 -07:00
  • fd71b11b1d move is_sm90_supported/is_sm100_supported to python/sglang/srt/utils.py (#9679) Lianmin Zheng 2025-08-27 03:34:29 -07:00
  • ae7428a8a7 fix mooncake store mla zero copy meta (#9678) huangtingwei 2025-08-27 15:43:16 +08:00
  • a3aee7c377 fix: HiRadixCache: fix prefetch completion race (#9397) Pablo Iyu Guerrero 2025-08-27 09:43:01 +02:00
  • 79e6a8a6ac support cuda 13.0 and trtllm kernel by Aug 25 2025 (#9495) Rain Jiang 2025-08-26 23:13:27 -07:00
  • 8f7b1c31e8 Add A100 fused MoE kernel configs for Dpsk (#9677) ehuaa 2025-08-27 11:49:48 +08:00
  • b9683be653 Support DeepSeek-V3.1 tool call (#9446) Xu Wenqing 2025-08-27 11:22:19 +08:00
  • a85363c199 [docs] Instructions for bench_serving.py (#9071) yhyang201 2025-08-27 09:30:57 +08:00
  • b21fdd5373 feat: (chat-template matching) enhance multimodal model detection with config.json (#9597) Kevin Tuan 2025-08-27 08:55:40 +08:00
  • c04c17edfa refactor(hicache): Introduce generic HiCacheStorageConfig for improved configuration management (#9555) hzh0425 2025-08-27 08:55:20 +08:00
  • 16a6d21b95 chore: enhance bench_serving for vlms with a new dataset of configurable image count and resolution (#9583) Mick 2025-08-27 08:42:54 +08:00
  • a530b3ffdc [RL] fix register the same ops multiple times (#9564) Stefan He 2025-08-26 16:24:44 -07:00
  • 603b3446dc Fix FA3 swa spec verify topk>1 (#9658) Ke Bao 2025-08-27 06:03:14 +08:00
  • b6c14ec0b4 add response_format support for completion API (#9665) cicirori 2025-08-27 00:01:29 +02:00
  • 43de1d7304 HiCache Storage fix host memory leak (#9648) Zhiqiang Xie 2025-08-26 10:49:40 -07:00
  • 79ce3688bb BugFix(hicache): Fix host indices out of bound error (#9637) hzh0425 2025-08-27 01:42:23 +08:00
  • 44ffe2cb72 Install py-spy by default for containers for easier debugging (#9649) fzyzcjy 2025-08-27 01:40:52 +08:00
  • 1a0896e9c0 [doc] add kimik2 --tool-call-parser (#9647) Xiaotong Jiang 2025-08-26 10:39:40 -07:00
  • 90313fb09a [router] add token bucket rate limiter (#9656) Chang Su 2025-08-26 10:36:26 -07:00
  • 3578eb1e9b [router] address worker load tracking consistency (#9523) Simo Lin 2025-08-26 06:40:51 -07:00
  • 0936c766ed Fix kimi k2 function calling format (#9606) Xiaotong Jiang 2025-08-26 00:50:59 -07:00
  • 0ef583b7de fix: allow user to specify function as role (#9635) GavinZhu-GMI 2025-08-26 15:47:20 +08:00
  • f7881a27f9 Add reasoning_effort param in TiktokenTokenizer.apply_chat_template (#9630) Liu Shaohui 2025-08-26 15:44:20 +08:00
  • fdff3167c5 [docs] Update README with additional highlights and resources for SGLang x AMD SF Meetup (#9640) Mingyi 2025-08-26 00:40:39 -07:00
  • cbc0e4d779 Fix lint for router (#9636) Stefan He 2025-08-26 00:38:53 -07:00
  • 4cd08dc592 model: Support nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 (#9301) Netanel Haber 2025-08-26 10:33:40 +03:00
  • f92b729d52 [new feat] ascend backend support fia fusion kernel (#8328) ZhengdQin 2025-08-26 14:13:08 +08:00
  • e2e378caba [router] add ut for mistral, llama, pythonic, and streaming tool parser (#9632) Simo Lin 2025-08-25 22:02:15 -07:00
  • dc1decc6af [router] add llama tool parser (#9629) Simo Lin 2025-08-25 20:43:36 -07:00
  • 03680f33be [router] add pythonic parser (#9628) Simo Lin 2025-08-25 20:40:06 -07:00
  • d4c5e53401 [router] add qwen tool parser (#9623) Simo Lin 2025-08-25 20:32:05 -07:00
  • 817c62a077 [router] add mistral tool parser (#9622) Simo Lin 2025-08-25 20:09:51 -07:00
  • 0ff7241995 Improve bench_one_batch_server script (#9608) Liangsheng Yin 2025-08-26 10:38:37 +08:00
  • 80dc76e11a [Fix] HiCache Bugfix & Mooncake Error Handling Enhance (#8901) ykwd 2025-08-26 10:05:10 +08:00
  • 9b08d975a0 [docs] Refactor, remove compiled results and add gpt-oss (#9613) Chayenne 2025-08-25 15:27:06 -07:00
  • a0a77d937b Fix Harmony reasoning parser for and auto-separation for gpt-oss models (#9190) Jonas 2025-08-26 00:26:26 +02:00
  • 24a8cee66d Fix GLM45v launch server cuda torch compile bug (#9554) Binyao Jiang 2025-08-25 13:46:28 -07:00
  • 3affa9dcc3 Fix GLM45 tool call multi-turn bug (#9500) Binyao Jiang 2025-08-25 13:46:13 -07:00
  • ea0696b924 [Performance] Batch Send from Tokenizer Manager. (#9436) Sundara Raman Ramachandran 2025-08-25 10:43:54 -07:00
  • 3aec3d4f8b [Doc] add LWS(LeaderWorkerSet) use case in sgl-router README (#9568) Bruce-x-1997 2025-08-25 23:32:31 +08:00
  • e3e97a120b chore: bump v0.5.1.post2 (#9592) Yineng Zhang 2025-08-25 03:45:09 -07:00
  • 051068673c chore: update config (#9591) Yineng Zhang 2025-08-25 03:41:09 -07:00
  • 9dcdf5da03 Tiny fix wrong comments (#9589) fzyzcjy 2025-08-25 18:08:10 +08:00
  • f8b757bcac fix: resolve tuning fused moe issue (#9587) Yineng Zhang 2025-08-25 01:41:15 -07:00
  • ebd9dbe71b fix: revert #8593 (#9581) Yineng Zhang 2025-08-25 01:29:06 -07:00
  • 938e986e15 chore: upgrade flashinfer 0.2.14.post1 (#9578) Yineng Zhang 2025-08-25 00:12:17 -07:00
  • 17d5eda887 bugfix for undefined logging functions in HarmonyBrowserTool & HarmonyPythonTool (#9229) Yuhao Zhou 2025-08-25 15:10:35 +08:00
  • 71a7f1d86f Offload tensors by sharding on GPU (#9536) fzyzcjy 2025-08-25 15:02:49 +08:00
  • 433266c125 Reintroduce memory usage fix (#9535) fzyzcjy 2025-08-25 15:02:31 +08:00
  • fda4792620 Update CUTLASS 4.2 & Enable K-Major Scale Factor for SM90 FP8 Blockwise Group GEMM (#9559) Qi Yuhang 2025-08-25 14:24:43 +08:00
  • a0b22f2f17 remove redundant rank0_log function. (#9560) miter 2025-08-25 14:17:55 +08:00
  • b5c6529e17 [PD] Improve disaggregation metrics output: update the metrics to keep reflecting real stats (#7317) SCDESPERTATE 2025-08-25 14:16:43 +08:00
  • ca4b86c564 fix: Update OpenAI client base URL in documentation (#9576) Xinyuan Tong 2025-08-25 14:06:57 +08:00
  • dd6ec02965 Add target module validation for init adapters (#9429) Beichen Ma 2025-08-24 20:24:50 -07:00