Commit Graph

  • 12b7a4fab0 [diffusion] performance: refactor diffusion fuse qkv and apply to qwen-image (#14793) Xiaoyu Zhang 2025-12-10 18:55:41 +08:00
  • 02f1e81e2d Revert "fix: checking if tokenizer is in cache before downloading from HF" (#14808) Yuhao Yang 2025-12-10 17:14:35 +08:00
  • 908c7186af [diffusion] CI: Add LoRA support to diffusion server configuration and test cases (#14697) Prozac614 2025-12-10 16:51:47 +08:00
  • 03836d85d2 [GLM-4.6V] Support Pipeline Parallelism for GLM-4.6V & GLM-4.1V (#14720) Yuan Luo 2025-12-10 16:40:12 +08:00
  • 87dbdddc93 [diffusion] profile: early exit when enough steps are captured to reduce the size of the trace file (#14803) Mick 2025-12-10 16:11:22 +08:00
  • 56e5c07424 fix b200 fa4 ci (#14788) b8zhong 2025-12-10 00:03:43 -08:00
  • 6c9c8da64d fix: add missing logic for SGLANG_USE_MODELSCOPE variable (#14794) yrk111222 2025-12-10 15:49:34 +08:00
  • 21028b5507 [RL] support weight reload for low-bit rollout (#9650) Peng Zhang 2025-12-10 15:44:01 +08:00
  • b0a25d0913 fix b200 ci (#14786) b8zhong 2025-12-09 23:08:41 -08:00
  • 793c98afaf handling incomplete rope_scaling config ci after transformers upgrade (#14784) Yuhao Yang 2025-12-10 14:56:16 +08:00
  • b1cbfce612 fix server args bug (#14725) TomerBN-Nvidia 2025-12-10 07:05:05 +02:00
  • b0f531ad28 Fix VLM accuracy thresholds for nightly tests (#14777) Alison Shao 2025-12-09 20:59:23 -08:00
  • 01835998e1 fix: race condition between validation and download locks (#14761) Alison Shao 2025-12-09 20:36:54 -08:00
  • 4285e99da7 [Auto Sync] Update data_parallel_controller.py, detokenizer... (20251209) (#14759) Lianmin Zheng 2025-12-09 18:38:38 -08:00
  • f077436831 [fix] Fix issues for in-flight weight updates (#14064) ShawnY112358 2025-12-10 10:12:19 +08:00
  • 5e8f544d1b Disable 8-gpu-b200 runner in PR tests (#14768) Alison Shao 2025-12-09 17:47:54 -08:00
  • c8d74feb06 fix: adding rate limit warning at verify token permission stage (#14756) Douglas Yang 2025-12-09 17:21:04 -08:00
  • cbc7dcdaa7 Re-add the API serving timing metrics. (#14744) Liangsheng Yin 2025-12-10 10:17:48 +09:00
  • a6dc7d2932 [ci]: Enable the new hf API (#14687) MingxuZh 2025-12-09 17:07:51 -08:00
  • 390406c46c [model-gateway] release gateway 0.2.4 (#14763) Simo Lin 2025-12-09 16:34:29 -08:00
  • 0c63fb9420 [Feature] Add LoRA support for embedding layers (#14177) Ethan (Yusheng) Su 2025-12-09 15:53:33 -08:00
  • 9ad02b799d [Perf] Optimize radix tree for cache-aware load balancin (#14758) Simo Lin 2025-12-09 15:35:29 -08:00
  • 18bd8e8d6d Improve CI by trying a warmup before unit tests (#14669) Lianmin Zheng 2025-12-09 15:17:59 -08:00
  • 036e64dafa move multi-item scoring functions in tokenizer manager into a separate file (#14740) Lianmin Zheng 2025-12-09 14:47:06 -08:00
  • 7c6fb3aa2d fix: making rate limit a warning instead of error (#14753) Douglas Yang 2025-12-09 13:41:09 -08:00
  • 8b0b6a45c8 fix: checking if tokenizer is in cache before downloading from HF (#14698) Douglas Yang 2025-12-09 13:08:07 -08:00
  • 55504df2f7 Add FP8 Blockwise GEMM Backend Flag --fp8-gemm-backend (#14379) b8zhong 2025-12-09 12:05:56 -08:00
  • 73df7a4e8d [SMG] perf: optimize tokenizer for reduced CPU and memory overhead (#14752) Simo Lin 2025-12-09 11:41:13 -08:00
  • 8b98bb768c [model-gateway] optimize core modules (#14751) Simo Lin 2025-12-09 10:54:07 -08:00
  • 15bc8cbd74 fix rope parameter initialization error caused by transformers v5.0 update (#14745) Yuhao Yang 2025-12-10 02:51:26 +08:00
  • 9496f12d00 [Model] Add PaddleOCR-VL Model Support (#12953) yudian0504 2025-12-10 02:16:02 +08:00
  • ab0048793c Revert "[Feat] Add received_time in serving_base" (#14743) Lianmin Zheng 2025-12-09 06:52:27 -08:00
  • 6ec77680df [diffusion] feat: support comparing batch perf (#14738) blahblah 2025-12-09 22:48:52 +08:00
  • 13680e5542 [Test] Skip STANDALONE speculative decoding tests for different hidden sizes (#14733) Alison Shao 2025-12-09 05:21:35 -08:00
  • 98c430e114 fix: prevent HugginqFace access when SGLANG_USE_MODELSCOPE is enabled (#12039) yrk111222 2025-12-09 21:02:05 +08:00
  • fe7f91ef82 [Feat] Add received_time in serving_base (#13432) zhanghaotong 2025-12-09 21:00:24 +08:00
  • cef5ba65b1 [Bugfix] Fix environ error in scheduler_runtime_checker_mixin.py (#14461) kun-llfl 2025-12-09 20:33:41 +08:00
  • 53d170883a Add fuse_marlin_moe test to ci and add new ep test (#14686) Xiaoyu Zhang 2025-12-09 20:17:38 +08:00
  • f0e948a0f1 fix the deepep 8 gpu unit test (#14601) Rain Jiang 2025-12-09 01:40:09 -08:00
  • 9a426fc5ef [CI] Move mistral large 3 basic to nightly (#14622) Alison Shao 2025-12-09 00:28:44 -08:00
  • 66772aa2b4 chore: add code owners for deepseek_v2.py (#14714) Yineng Zhang 2025-12-08 23:24:06 -08:00
  • b626334475 fix: make override DeepseekV2Model work (#14707) Yineng Zhang 2025-12-08 23:11:18 -08:00
  • 0f8bd55f3e [CI] Fix Llama 3.1 8B FP4 CI (#14699) b8zhong 2025-12-08 22:27:15 -08:00
  • da3dc497b0 Fix dp-aware incompatible with completions and chat completions APIs (#14647) fzyzcjy 2025-12-09 13:24:53 +08:00
  • 817daba062 Tiny extract select_worker_min_load (#14648) fzyzcjy 2025-12-09 13:22:10 +08:00
  • af60cad05d [ci][smg] fix docker release ci and add it to pr test (#14683) Simo Lin 2025-12-08 20:56:38 -08:00
  • ce4e836be5 Add per-request decode tp size (#14678) Lianmin Zheng 2025-12-08 20:24:31 -08:00
  • 0e0b0c0566 Revert "[Bug] fix not desired disable fused share experts caused by r… (#14676) Yineng Zhang 2025-12-08 20:06:52 -08:00
  • e6f0ddda44 [CI] Migrate Eagle 1-GPU tests to test/registered/ (#14529) Alison Shao 2025-12-08 19:56:36 -08:00
  • af20657cd4 Tiny support sgl-router http response status code metrics (#14689) fzyzcjy 2025-12-09 11:54:35 +08:00
  • 08da4c2618 [Bugfix] Fix KeyError for Mistral-Large-3 rope_scaling config (#14627) Alison Shao 2025-12-08 19:16:00 -08:00
  • ef3f8c97e1 Add ffmpeg into sglang docker - required by transformers multimodal V… (#14679) Binyao Jiang 2025-12-08 18:00:23 -08:00
  • e5201bda34 [CI] Unblock gb200 cutedsl test (#14469) Baizhou Zhang 2025-12-08 17:58:25 -08:00
  • 60d36e7be7 [NPU] chore: bump basic software version to 8.3.rc2 (#14614) Even Zhou 2025-12-09 09:14:27 +08:00
  • 6f657070ef [SMG]feat: implement TokenGuardBody for managing token return (#14653) Jimmy 2025-12-09 08:44:46 +08:00
  • c106b54b57 Aiter fp8 kv cache (#13147) kk 2025-12-09 08:39:53 +08:00
  • 119fd956fb Tiny support printing requests in bench_serving for observability (#14652) fzyzcjy 2025-12-09 08:27:58 +08:00
  • eac5b66485 [RadixTree] Optimize the Time Complexity of Node Retrieval Operation from O(n*m) to O(n) (#13334) PiteXChen 2025-12-09 07:55:09 +08:00
  • 07404d7689 [HiCache] fix condition check when use decode offload (#14489) Francis 2025-12-09 07:52:40 +08:00
  • 93043f7b13 [CI] Increase max-parallel to 15 for high priority PRs (#14675) Alison Shao 2025-12-08 15:40:37 -08:00
  • edde5e5d40 [model-gateway] add OTEL integration to grpc router (#14671) Simo Lin 2025-12-08 14:45:05 -08:00
  • 6abb8051e8 Bump up diffusers to latest official release version (#14670) Binyao Jiang 2025-12-08 13:41:01 -08:00
  • 2e3946d889 Fix cache-aware router should pick min load instead of min tenant size (#14650) fzyzcjy 2025-12-09 05:33:34 +08:00
  • 32f8b6064e improve default glm mtp setting (#14457) b8zhong 2025-12-08 13:27:13 -08:00
  • b9bef31a15 fix: use .get() when accessing strict mem-check env variable (#14657) Yuhao Yang 2025-12-09 05:25:42 +08:00
  • 8550822d6b [model-gateway] Optimize memory usage in HTTP router (#14667) Simo Lin 2025-12-08 13:10:31 -08:00
  • 8810152e88 vlm: Use fa3 as the default backend for qwen3 vl (#14634) Mick 2025-12-09 04:56:20 +08:00
  • 7bf16c6339 [model-gateway] fix WASM arbitrary file read security vol (#14664) Simo Lin 2025-12-08 12:11:16 -08:00
  • 39f9a9c2a5 [model-gateway] reduce cpu overhead in grpc router (#14663) Simo Lin 2025-12-08 11:54:56 -08:00
  • d69ecc19b8 [model-gateway] reducing cpu overhead in various of places (#14658) Simo Lin 2025-12-08 09:44:40 -08:00
  • 763888b5a8 [AMD] change fused rms quant interface for aiter upgrade (#14497) yctseng0211 2025-12-09 01:09:23 +08:00
  • 9a327bdfcf chore: bump SGLang version to 0.5.6.post1 (#14651) sglang-bot 2025-12-08 08:35:28 -08:00
  • 2de98010b5 chore: bump sgl-kernel version to 0.3.19 (#14649) sglang-bot 2025-12-08 06:53:08 -08:00
  • 8200fb56cb update transformers package version to 5.0.0rc0 (#14356) Yuhao Yang 2025-12-08 22:46:01 +08:00
  • cb4cdb43a4 Fix dp-aware incompatible with service-discovery (#14629) fzyzcjy 2025-12-08 22:39:27 +08:00
  • 80cfca50bc [diffusion] chore: further refine output resolution adjustment logic (#14558) Mick 2025-12-08 19:08:38 +08:00
  • 7871593cc8 [cpu] Implement all gather/reduce for arm64 cpu (#12527) Yibo Cai 2025-12-08 19:03:04 +08:00
  • 4a62a0e3cd chore: bump sgl-kernel version to 0.3.19 (#14632) sglang-bot 2025-12-08 03:02:24 -08:00
  • 12a08efc20 [diffusion] feat: add support for LoRA layers in transformer_2 within LoRAPipeline (#14606) Prozac614 2025-12-08 17:57:13 +08:00
  • 06836ad02a [Reasoning + Structured Output] make reasoning compatible with structured output (#12551) Muqi Li 2025-12-08 17:28:39 +08:00
  • f72a77038f modify the sgl-kernel to be compatible with transformers 5.x. (#14625) Yuhao Yang 2025-12-08 16:39:00 +08:00
  • aeff0d386b Fix amd rope definition (#14556) Qiaolin Yu 2025-12-07 23:47:03 -08:00
  • cf0478d602 [Glm46v] Bug fix for accuracy drop and unable to launch server (#14585) Binyao Jiang 2025-12-07 23:45:02 -08:00
  • a2ca9bd4f1 Super tiny fix unused code in router (#14618) fzyzcjy 2025-12-08 15:12:09 +08:00
  • 36361adcbf [DLLM] Add initial cuda graph support (#14203) Tiwei Bie 2025-12-08 14:12:35 +08:00
  • 661e9775d0 [2/2] Add rope kernel in sgl-kernel (#14452) Qiaolin Yu 2025-12-07 21:37:29 -08:00
  • 2970f22917 [model-gateway] fix WASM unbounded request/response body read vuln (#14612) Simo Lin 2025-12-07 20:47:25 -08:00
  • c08b780fe0 Super tiny remove unused select_worker_pair (#14609) fzyzcjy 2025-12-08 12:32:38 +08:00
  • 8fbf7dd56f [model-gateway] refactor otel to be more efficient (#14604) Simo Lin 2025-12-07 20:09:25 -08:00
  • 1915a1f8a7 Super tiny remove unneeded policy flag (#14608) fzyzcjy 2025-12-08 11:35:07 +08:00
  • 85d0ccfac0 Tiny fix missing policy decision recording (#14605) fzyzcjy 2025-12-08 11:24:27 +08:00
  • a4ffd665c0 [model-gateway] fix WASM memory limit per module (#14600) Simo Lin 2025-12-07 19:15:19 -08:00
  • 559202b544 [CI] Fix unit-test-backend-8-gpu-b200 running on every /rerun-stage (#14591) Alison Shao 2025-12-07 18:42:04 -08:00
  • f57d4fe78e [feat] use cachebuffer to store mm feature to speedup hash (#14386) Nicholas 2025-12-08 10:35:20 +08:00
  • b7b7524e95 [Tool Call] Fix DeepSeekV32Detector skipping functions with no params in streaming mode (#14573) wentx 2025-12-08 10:32:10 +08:00
  • 6799847ebf [CI]Unblock and split spec v2+dp test (#14551) Baizhou Zhang 2025-12-07 17:39:25 -08:00
  • 03b835e7d1 Refactor tuning block wise kernel and opt Qwen/Qwen3-VL-32B-Instruct-FP8 (#14141) Xiaoyu Zhang 2025-12-08 09:24:58 +08:00
  • aff1238ef2 [model-gateway] reorganize metrics, logging, and otel to its own module (#14590) Simo Lin 2025-12-07 16:50:39 -08:00
  • 5e2cda6158 [model-gateway] Fixed WASM Security Vulnerability - Execution Timeout (#14588) Simo Lin 2025-12-07 16:05:53 -08:00
  • b0bbc7f53d [model-gateway] extra accumulator and tool handler in oai router (#14587) Simo Lin 2025-12-07 15:52:12 -08:00