Commit Graph

  • 385a35bd11 [AMD][MORI] Fix MTP crash with FP4/FP8 dispatch and add NEXTN dispatch env vars. (#20647) Duyi-Wang 2026-03-17 16:13:42 +08:00
  • ee106757df [diffusion] fix: fix Diffusers backend ignores model-specific sampling parameter (#20080) Junhao Liu 2026-03-17 01:10:46 -07:00
  • 9a697ceabb [Fix #20389] Illegal memory access in triton attention for large token counts (#20390) akhilg-nv 2026-03-17 00:42:11 -07:00
  • e3277b3be2 [diffusion]: remove stale offload-manager in LTX2 AV denoising (#20624) Ratish P 2026-03-17 12:44:00 +05:30
  • 18dd5c972d [AMD] Fix AMD CI : stage-c-large-8-gpu-amd-mi35x (#20521) YC Yen-Ching Tseng 2026-03-17 14:53:33 +08:00
  • 025691cd9e [diffusion] chore: bump up cache-dit & support quant for diffusers backend (#20361) DefTruth 2026-03-17 12:51:31 +08:00
  • 079a1fd35e [Bugfix] Fix write-through events not processed when scheduler is idle (#20560) Rocky Song 2026-03-17 13:49:59 +09:00
  • 5d5c31c6e4 [PP] Add CP pyobj broadcasting when enable dynamic CPP (#20738) Shangming Cai 2026-03-17 12:20:11 +08:00
  • 855ec7017d Add check to provide hicache-storage-backend when enabling kv caching on Decode Side in PD Disaggregation (#20732) MMuzzammil1 2026-03-17 12:25:14 +09:00
  • 943f34f642 Add NCCL/RCCL pre-warming to reduce P99 TTFT cold-start latency (#20477) Hubert Lu 2026-03-16 20:23:14 -07:00
  • 5ec49a5309 Revert "[diffusion] CI: use dedicated HF token for accessing restricted models" (#20737) Mick 2026-03-17 11:22:20 +08:00
  • 210d0fbaef Update ascend docs (#20674) amote-i 2026-03-17 11:14:26 +08:00
  • e4d06b3db2 Fix /generate JSON serialization for non-finite top_logprobs (#20714) Jay Shaik 2026-03-17 08:37:12 +05:30
  • 515b3a323d feat: support human-readable suffixes (25.6k, 1M, 1Mi) for token CLI (#20577) shuwenn 2026-03-17 11:05:33 +08:00
  • 3cafb02e5b AMD/ROCm: update daily image repo to lmsysorg/sglang-rocm (#20720) HAI 2026-03-16 20:04:52 -07:00
  • 9f56b471aa [Network] Use NetworkAddress for dist_init_method and loopback fallbacks (#20657) psaab 2026-03-16 19:59:49 -07:00
  • 71a54c1c42 update CODEOWNERS (#20733) Qiaolin Yu 2026-03-16 19:41:26 -07:00
  • 4dbec2dd2b [typo] Fix typos in comments and log messages in common.py (#20723) Jason Yao 2026-03-17 10:26:59 +08:00
  • 7d87a6a071 Fix spec v1 token_ids_logprobs (#20718) Qiaolin Yu 2026-03-16 19:23:28 -07:00
  • 474a851ae3 [diffusion] fix: fix sampling params incorrectly override in cli (#20689) Mick 2026-03-17 08:48:10 +08:00
  • 1eea744855 [diffusion] CI: enable UT (#20690) Mick 2026-03-17 07:44:04 +08:00
  • 2ccdb7373e [diffusion] CI: fix consistency test workflow (#20704) Yuhao Yang 2026-03-17 07:42:30 +08:00
  • 826eb21bca Fix sglang-kernel dependency on CI runners (#20715) Baizhou Zhang 2026-03-16 15:23:01 -07:00
  • 5ef5806160 [Nemotron] Small reasoning parser fix (#20284) roikoren755 2026-03-16 22:29:40 +02:00
  • 70a6fb53af Enable embedding lookup/lora_a logic for chunked backend (#17692) Bruce Wu 2026-03-16 11:37:58 -07:00
  • 061ec582bf fix: adding teacache.params back to sampling params as intended (#20665) Douglas Yang 2026-03-16 11:27:06 -07:00
  • a4528a5737 [CI] Bring back CI test for Mamba PD Disaggregation (#20675) Shangming Cai 2026-03-17 01:04:21 +08:00
  • 289cbcf482 fix: support PP2+CP8+TP8 (PP with context parallelism) (#19548) ybyang 2026-03-17 00:51:47 +08:00
  • 6489f77733 [Diffusion] Fix compile graph broken by flashinfer rope (#20699) Xiaoyu Zhang 2026-03-16 23:14:27 +08:00
  • d3c0f4376a Fix AssertionError crash in disagg prefill inflight queue with PP (#20686) Du Bin 2026-03-16 22:38:59 +08:00
  • 15097c5c3b Release sglang kernel 0.4.0 (#20440) Xiaoyu Zhang 2026-03-16 20:34:58 +08:00
  • 3d58cd16d9 [DP Attention] Optimize dp_padding_mode selection for dp_size=1 in extend mode (#20406) sky 2026-03-16 18:44:42 +08:00
  • 549fbcc864 [5/N] (Elastic EP) Use GPU P2P to exchange expert weights during EPLB as much as possible (#12068) Xun Sun 2026-03-16 18:40:58 +08:00
  • 3055b6906d [Diffusion] Document torch.compile graph-break checks in diffusion benchmark skills (#20681) Xiaoyu Zhang 2026-03-16 17:41:40 +08:00
  • 485597e651 [diffusion] fix: fix some sampling args passed via cli are omitted (#20630) Mick 2026-03-16 16:55:30 +08:00
  • 895e56097c Add NPU basic function testcases (#19382) Sugar920 2026-03-16 15:09:56 +08:00
  • e96a3752a0 Update timeout for some ci tests (#20664) Baizhou Zhang 2026-03-15 23:58:30 -07:00
  • 42f18fe560 [HiCache] fix: release write-through lock_ref during decode (#20049) shuwenn 2026-03-16 14:49:31 +08:00
  • 39336f5812 Precompute swa cache location (#20449) Ke Bao 2026-03-16 14:38:08 +08:00
  • 135af6dc92 [EPD][VLM] support video/audio input (#17824) Zheng Wengang 2026-03-16 14:18:21 +08:00
  • 738cbde902 [PD] Make pending reqs resolving more robust (#20505) Shangming Cai 2026-03-16 14:12:13 +08:00
  • a92872cbc3 [AMD] CI - skip tests on mi325 (#20660) YC Yen-Ching Tseng 2026-03-16 13:35:04 +08:00
  • 223a4e485b [test] Fix bfloat16 tolerance in test_selective_state_update (#20345) Roopak Srivastava 2026-03-16 09:32:01 +05:30
  • 97b2a89334 [RadixTree][8/N Refactor]: unify lock interface (#20330) pansicheng 2026-03-16 11:49:51 +08:00
  • f0458e0b49 [Utils] Move network/socket utilities from common.py to network.py (#20646) Liangsheng Yin 2026-03-15 20:35:24 -07:00
  • afc71bae3a feat: Add 'none' reasoning effort to ChatCompletionRequest (#20556) Javier Torres 2026-03-15 20:25:48 -07:00
  • da1793f63a update ascend feature docs (#20506) amote-i 2026-03-16 11:09:20 +08:00
  • f4393bf3f6 Fix correctness test issue for bench_one_batch (#20650) gaopengff 2026-03-16 11:05:36 +08:00
  • e1eb25880f [Diffusion] Add a benchmark for rmsnorm/fuse_add_rmsnorm (#20632) Xiaoyu Zhang 2026-03-16 09:50:33 +08:00
  • e9fae69e5f docs: align environment variable reference with environ defaults (#20419) AnonTokyo 2026-03-16 09:07:29 +08:00
  • 3879c466b4 [CI] Add Nemotron 3 Super 120B nightly 8-GPU tests (#20616) Mohammad Miadh Angkad 2026-03-16 09:03:20 +08:00
  • 35c249b4de [OpenAI] Log raw request payload for --log-requests (#20605) Zhirui 2026-03-16 08:45:00 +08:00
  • d852f26cb6 Fix dual-stack socket handling: IPV6_V6ONLY, IPv4-first, is_port_available all-family check (#20643) Liangsheng Yin 2026-03-15 17:17:23 -07:00
  • 7c498a6538 [DOC] add documents for encoder global mm cache (#20636) Teng Ma 2026-03-16 07:44:21 +08:00
  • c2245d9b55 ci: add runner dropdown to rerun-ut workflow dispatch (#20645) Liangsheng Yin 2026-03-15 16:40:14 -07:00
  • 53f831691a fix: propagate grammar errors and improve llguidance backend (#20467) jellysnack 2026-03-16 02:11:18 +03:00
  • 116aef8504 [Test] Move embedding tests into test/registered/embedding/ and unit/ (#20642) Liangsheng Yin 2026-03-15 14:48:43 -07:00
  • 1145805e7d Fix socket utilities and reserve_port for IPv6 dual-stack support (#20491) psaab 2026-03-15 14:29:10 -07:00
  • 8397a8a301 [CI] Extract wait-for-jobs composite action and gate on stage-a-cpu-only (#20641) Liangsheng Yin 2026-03-15 13:52:48 -07:00
  • 23c191afb6 fix(docs): correct quantization documentation (#20301) (#20619) Mook 2026-03-15 09:33:12 -07:00
  • c3483e8e97 [CI] Move existing unit tests into unit directory (#20631) Ke Bao 2026-03-15 23:25:18 +08:00
  • e2be31824f [CI] Add ut coverage tool (#20628) Ke Bao 2026-03-15 21:13:45 +08:00
  • 45dd06f4e0 Remove insecure auto-format label workflow (#20629) Xiaoyu Zhang 2026-03-15 21:04:29 +08:00
  • 0cff070828 [CI] Add READMEs for unit test directory structure (#20626) Ke Bao 2026-03-15 20:27:05 +08:00
  • 7142a594f9 [CI]: Add CI For HiMambaRadixTree and qwen3.5 (#20540) hzh0425 2026-03-15 20:16:12 +08:00
  • 1c456a0af5 VLM: add Conv2dLayer/Conv3dLayer to fix PyTorch 2.9.1 CuDNN Conv3d (#20282) Yuhao Yang 2026-03-15 19:17:44 +08:00
  • f07529b947 [diffusion] CI: use dedicated HF token for accessing restricted models (#20620) Mick 2026-03-15 19:15:28 +08:00
  • 4f91b3069e Skip flaky test in CI for disaggregation decode offload (#20622) Shangming Cai 2026-03-15 16:30:34 +08:00
  • 7c773ddb0a [Fix] Slice input_embeds to extend_input_len in prepare_for_extend (#20376) Kit Fraser-Taliente 2026-03-15 00:07:05 -07:00
  • 7458407437 Fix InternVL and vision attention for non-CUDA backends (e.g. XPU) (#19997) Juan Muneton 2026-03-14 23:24:41 -07:00
  • 1ac6a26464 fix: Nemotron chunk size alias (#20458) shuwenn 2026-03-15 14:23:39 +08:00
  • fc7f9c1de7 Rename --stream-output to --incremental-streaming-output (#20614) Liangsheng Yin 2026-03-14 23:22:33 -07:00
  • 538acb4c46 fix: Add .text property to HttpResponse to prevent AttributeError (#20518) shuwenn 2026-03-15 13:59:32 +08:00
  • a6ecf050be diffusion: fix helios accuracy issue (#20036) Yuhao Yang 2026-03-15 13:55:51 +08:00
  • 6c5bf53a36 [Doc] Clarify that --chat-template is required for Qwen3-Reranker (#20596) Matt Van Horn 2026-03-14 16:43:48 -07:00
  • 93afe15b43 chore: bump flashinfer version to 0.6.6 (#20480) sglang-bot 2026-03-14 13:05:10 -07:00
  • 574dbe23b2 Add piecewise cuda graph for Qwen3-Next FP8 flashinfer_trtllm moe backend (#18184) Xiaowei Wang 2026-03-15 04:03:31 +08:00
  • 3e643967e6 [CI] Add Nemotron 3 Super 120B CI tests for BF16 and NVFP4 (#20575) Mohammad Miadh Angkad 2026-03-15 03:30:27 +08:00
  • 39008955ff Revert "[AMD][MORI] Fix MTP crash with FP4/FP8 dispatch and add NEXTN dispatch env vars." (#20602) Baizhou Zhang 2026-03-14 12:12:42 -07:00
  • 5ab2cfe9a8 [Diffusion] Clean upstream fa3 in hopper (#20576) Xiaoyu Zhang 2026-03-14 23:41:23 +08:00
  • 22e67876d6 [Omni] Optimize AudioEncoder for Qwen3_Omni_Thinker (#18185) Yuan Luo 2026-03-14 23:00:17 +08:00
  • 574aa2d723 [diffusion]: remove stale offload-manager cleanup in denoising stage (#20587) Ratish P 2026-03-14 20:26:57 +05:30
  • 25e38216b6 [kernel slimming] Clean many useless sgl-kernel deprecated kernels (#20277) Xiaoyu Zhang 2026-03-14 16:45:54 +08:00
  • 75a7879fd4 [Model] Support Nemotron 3 Super NVFP4 (#20407) Mohammad Miadh Angkad 2026-03-14 15:56:26 +08:00
  • c95dc88f86 [CI] migrate ascend-gptq from test/srt to test/registered (#19628) SoluMilken 2026-03-14 15:28:57 +08:00
  • f9e4221b71 [Diffusion] add mova and hunyuanvideo to perf skills (#20563) Xiaoyu Zhang 2026-03-14 13:49:50 +08:00
  • 99a3b25c9b [PP] Fix recv tensor dict potential race condition (#20341) Shangming Cai 2026-03-14 13:35:01 +08:00
  • c330b687a1 [Bugfix] Fix GLM-4.6V vision regression in glm4v_moe and glm_ocr (#20463) Xinyuan Tong 2026-03-14 04:48:28 +00:00
  • dfd0a77a9a [bugfix] Add prev_prefix_len parameter to HiMambaRadixCache's _insert_helper() (#20539) ziruiliu 2026-03-14 09:54:14 +08:00
  • 4659b08fcf update CI_PERMISSIONS.json (#20551) Qiaolin Yu 2026-03-13 14:52:06 -07:00
  • 0eea80bc00 [AMD][MORI] Fix MTP crash with FP4/FP8 dispatch and add NEXTN dispatch env vars. (#20453) Duyi-Wang 2026-03-14 05:03:17 +08:00
  • c37ef7f18b [AMD] diffusion refactor: move ROCM VAE optimization to Platform abstraction (#20496) YC Tseng 2026-03-14 04:10:05 +08:00
  • d764f414a1 [AMD] fix mori unittest (#20524) billishyahao 2026-03-14 00:37:45 +08:00
  • a16f03c1e3 Revert "feat: add banner to sgl-model-gateway (#20471)" (#20536) Simo Lin 2026-03-13 09:20:32 -07:00
  • 654fc02cf1 [gRPC] Extract gRPC servicer into standalone package (#20478) Simo Lin 2026-03-13 09:13:29 -07:00
  • be7a0311a0 [Diffusion] Fix and validate diffusion skills benchmarking/profiling workflow (#20528) Xiaoyu Zhang 2026-03-13 21:11:37 +08:00
  • b1246c50f8 Fix chunked prefill and KV cache leaks for streaming sessions (#20476) Leon Gao 2026-03-13 02:36:55 -07:00
  • 287dc12b05 Fix hicache log metrics (#20504) Ke Bao 2026-03-13 16:29:58 +08:00
  • f8668d9e78 [Fix] Add fallback for flashinfer allreduce fusion (#20384) Baizhou Zhang 2026-03-13 01:24:55 -07:00
  • b638b25b22 [diffusion] UX: suppress excessive logging from httpx and httpcore (#20452) Mick 2026-03-13 14:43:09 +08:00