Commit Graph

  • 7f35c46efb Tiny add sglang:http_requests_active metric (#16479) fzyzcjy 2026-01-05 17:20:31 +08:00
  • da2f8cc33f [minor] reduce 1 unnecessary add (#16474) DarkSharpness 2026-01-05 17:00:56 +08:00
  • 4308c25b47 Super tiny move test folder (#16447) fzyzcjy 2026-01-05 16:29:26 +08:00
  • 4d737db857 Tiny fix prefill delayer not support non-fcfs schedule policy (#16471) fzyzcjy 2026-01-05 16:28:54 +08:00
  • 9d6029fb92 Fix TokenizerManager bottleneck for offline generation (#16456) fzyzcjy 2026-01-05 16:28:25 +08:00
  • dcfb92ddbc [diffusion] fix: fix using the sub-optimal fa kernel (#16382) Qi Yuhang 2026-01-05 16:25:41 +08:00
  • bebd625ba1 EVS Framework: Support NemotronH_Nano_VL_V2 (#14051) Netanel Haber 2026-01-05 10:18:07 +02:00
  • b12258bfaa [model-gateway] Optimize HTTP Router Fan-out: Replace Serial Execution with Concurrent Streams (#16042) Praneth Paruchuri 2026-01-05 13:08:45 +05:30
  • 012dc5866d [Auto Sync] Update scheduler.py (20260104) (#16424) Lianmin Zheng 2026-01-04 23:32:56 -08:00
  • 7f6a678f8f [model-gateway] add GPU allocator and model pool infrastructure for parallel E2E tests (#16460) Simo Lin 2026-01-04 23:30:32 -08:00
  • 399ca037b1 [bug] fix lint blocking ci (#16464) Simo Lin 2026-01-04 22:57:33 -08:00
  • d93f37a625 Add harvenstar to CI_PERMISSIONS.json Kangyan-Zhou 2026-01-04 22:35:26 -08:00
  • e6fe092dcc [model-gateway] rename py_test to e2e_test (#16454) Simo Lin 2026-01-04 22:24:34 -08:00
  • f02d82211a [model-gateway]: move PD configuration conflict checks to model gateway (#16088) Ratish P 2026-01-05 10:16:07 +04:00
  • 10174e1114 Revert "[grpc] update api to scheduler in grpc request manager" (#16387) Yuhao Yang 2026-01-05 14:05:39 +08:00
  • 2138ff48c6 Revert "[FEAT] optimize tensor zmq transfer for multimodal inputs" (#16386) Yuhao Yang 2026-01-05 14:05:26 +08:00
  • e53160bb31 [model-gateway] reorganize integration tests into logical subdirectories (#16451) Simo Lin 2026-01-04 21:49:20 -08:00
  • 4ea6a11c83 [CI] Fail wheel build when sgl-kernel artifacts are missing (#16450) Xiaoyu Zhang 2026-01-05 13:46:40 +08:00
  • 2181bc9e51 [model-gateway] Delete Python integration_mock tests (#16448) Simo Lin 2026-01-04 21:34:06 -08:00
  • 520c048d55 [diffusion] CI: add script for automatically generation ci perf baseline (#16389) Xiaoyu Zhang 2026-01-05 13:18:35 +08:00
  • 561a3e04f6 remove redundant mamba_cache clearing actions (#16180) yudian0504 2026-01-05 13:02:07 +08:00
  • f84487af59 [model-gateway] : Rust integration tests for integration_mock replacement (#16441) Simo Lin 2026-01-04 21:01:47 -08:00
  • 078270473a [Doc] Default lora backend: csgmv (#16444) Huapeng Zhou 2026-01-04 20:45:49 -08:00
  • c63e9cb29e Super tiny fix CI (#16437) fzyzcjy 2026-01-05 12:34:49 +08:00
  • 12cde0df99 [SPEC_V2] Fix Acclen drop when enabling DP Attention for Spec-Overlap (#16310) YAMY 2026-01-04 19:39:23 -08:00
  • a7fd810842 Allow editable install without .git with add fallback version in pyproject.toml (#16435) Liangsheng Yin 2026-01-05 11:17:20 +08:00
  • ca80c19b55 Revert "[diffusion] feat: support warmup with resolutions" (#16433) Kangyan-Zhou 2026-01-04 18:44:05 -08:00
  • 1e7b326482 Super tiny fix main code (#16432) fzyzcjy 2026-01-05 10:34:25 +08:00
  • e267ca0beb [diffusion] doc: document LoRA support in CLI (#16375) CHEN Xi 2026-01-05 10:20:30 +08:00
  • 9a8ba3c189 [diffusion] feat: support warmup with resolutions (#16330) Mick 2026-01-05 10:16:26 +08:00
  • 0fee6bc632 [JIT kernel] Apply jit per_tensor_quant_fp8 kernel (#15836) Xiaoyu Zhang 2026-01-05 10:15:00 +08:00
  • 0ff3747ca1 [model-gateway]: move unit tests to bindings/python/tests/ (#16430) Simo Lin 2026-01-04 18:10:35 -08:00
  • f8411ded6e ci: migrate 1-GPU model tests to test/registered/models/ (#16414) Alison Shao 2026-01-04 18:08:01 -08:00
  • 249c356331 Super tiny update tokenizer benchmark (#16429) fzyzcjy 2026-01-05 09:14:52 +08:00
  • 12df16607b Tiny speed up kimi detokenizer by 10x (#16427) fzyzcjy 2026-01-05 09:12:05 +08:00
  • 55d112dc79 fix: enable multi-threading for h200 tests (#16413) Douglas Yang 2026-01-04 14:43:53 -08:00
  • a1ed247fd7 feature: add runner online count to failure monitor (#16408) Douglas Yang 2026-01-04 13:24:04 -08:00
  • 87699d48eb fix: only publish trace from tp 0 (#16411) Douglas Yang 2026-01-04 12:17:08 -08:00
  • 52c604342c chore: print test list at beginning and end of run_suite.py (#16334) Alison Shao 2026-01-04 11:52:52 -08:00
  • ff0f370f85 ci: migrate MoE tests to test/registered/moe/ (#16127) Alison Shao 2026-01-04 11:51:12 -08:00
  • 26f9e20755 ci: migrate quantization kernel tests to test/registered/quant/ (#16323) Alison Shao 2026-01-04 11:49:29 -08:00
  • 828cd8936f Introduce sgl-kernel Dockerfile (#14066) Yingchun Lai 2026-01-05 03:19:08 +08:00
  • 4436dc0f6c [model-gateway] improve lock contention and allocation in middleware (#16405) Simo Lin 2026-01-04 10:40:07 -08:00
  • cf6800f6cc [model-gateway] optimize Vec and HashMap allocations in responses api (#16406) Simo Lin 2026-01-04 10:07:55 -08:00
  • f16606d6f6 Super tiny code cleanup (#16401) fzyzcjy 2026-01-04 23:00:09 +08:00
  • 387fad2f74 Tiny add detokenization benchmarks (#16400) fzyzcjy 2026-01-04 22:53:38 +08:00
  • 76bc07a335 Move swa memory pool to a seperate file (#16347) Ke Bao 2026-01-04 22:39:30 +08:00
  • b328cd20bb Skip local attn init metadata for mimo swa model (#16349) Ke Bao 2026-01-04 22:38:36 +08:00
  • bf32cd8397 [npu] update model and feature supported for ascend npu (#16390) Hexq0210 2026-01-04 21:50:20 +08:00
  • 27a0830593 Tiny add manual policy benchmark to CI (#16392) fzyzcjy 2026-01-04 19:10:35 +08:00
  • 53479e226b Tiny refactor router benchmark github actions (#16374) fzyzcjy 2026-01-04 18:00:12 +08:00
  • ff978e7d4c Update readme (#16391) Lianmin Zheng 2026-01-04 01:57:54 -08:00
  • 16e00651dd fix: health check adding test delays (#16381) Douglas Yang 2026-01-04 00:28:32 -08:00
  • 1b2b95d8d1 Super tiny cleanup unused function (#16376) fzyzcjy 2026-01-04 15:08:12 +08:00
  • 9cac3c8628 Add liusy58 into CI_PERMISSION (#16373) Shangming Cai 2026-01-04 13:36:35 +08:00
  • 8d58b3dcdd [model-gateway] Clear architectual debt in responses API (#16359) Chang Su 2026-01-03 21:35:28 -08:00
  • 5f3eb377e0 [VLM] Support request level max_dynamic_patch for OpenAI request (#16268) Yuan Luo 2026-01-04 13:04:43 +08:00
  • 229938805f Make personal configs optional in SGLang's official docker image. (#16365) Liangsheng Yin 2026-01-04 12:27:13 +08:00
  • d7aa0ce72f Support multiple engines for router simulation in schedule simulator (#16368) fzyzcjy 2026-01-04 11:46:05 +08:00
  • e797f0c570 Support offline generation scenario for prefill delayer (#16363) fzyzcjy 2026-01-04 11:05:10 +08:00
  • 5d4b7c78bf Fix memory leak in prefill delayer (#16358) fzyzcjy 2026-01-04 10:11:18 +08:00
  • 9bd64d739b [VLM] Add doc for ViT CUDA Graph (#16343) Yuan Luo 2026-01-04 10:05:23 +08:00
  • 24e116ef6a Enhance schedule simulator and add several e2e scenarios (#16357) fzyzcjy 2026-01-04 10:02:38 +08:00
  • 8ca9597057 Tiny support sticky routing algorithm in schedule simulator (#16355) fzyzcjy 2026-01-04 08:36:52 +08:00
  • 216ea910ad Tiny support forwarding SMG routing key to engine for dumping and avoid string allocation (#16352) fzyzcjy 2026-01-04 08:33:35 +08:00
  • f4ab2ec5be Add unified metrics collection framework (v1) (#16064) Hudson Xing 2026-01-04 08:30:38 +08:00
  • 25fa2ac290 Convert cu_seqlens to CPU for npu_flash_attention_unpad operator (#15434) fy 2026-01-04 08:16:42 +08:00
  • ef5ac6f01e Support shared prefix (gsp) in schedule simulator (#16353) fzyzcjy 2026-01-04 08:03:16 +08:00
  • c88aaf22c7 Support in-flight request age metrics for router (#16341) fzyzcjy 2026-01-04 07:38:53 +08:00
  • e139d2aa76 Super tiny fix CI (#16351) fzyzcjy 2026-01-04 07:24:46 +08:00
  • 2337b1bbb0 Refactor and fix prefill delayer (scheduler enhancer) (#16269) fzyzcjy 2026-01-04 07:15:06 +08:00
  • 7bc13c906c [grpc] update api to scheduler in grpc request manager (#16350) Chang Su 2026-01-03 10:54:31 -08:00
  • 877c8e3a96 Tiny refactor router test contexts (#16340) fzyzcjy 2026-01-04 01:44:10 +08:00
  • 66dfb8c156 Tiny fix non-PD router http header missing whitelist (#16339) fzyzcjy 2026-01-04 01:43:02 +08:00
  • dcacc492d0 [diffusion] fix: fix RuntimeError in SageAttention3 on Blackwell with Qwen-Image (#16335) Chi McIsaac 2026-01-03 09:12:53 -05:00
  • 87ef05e2e1 Support simple schedule simulator (#16344) fzyzcjy 2026-01-03 22:05:40 +08:00
  • b65c9889a0 Fix incorrect running batch size in prefill stats (#15941) fzyzcjy 2026-01-03 20:49:54 +08:00
  • c7c0d97fc6 Tiny support whitelisted headers in request logging (#16342) fzyzcjy 2026-01-03 20:10:58 +08:00
  • 6bc5a52fd2 [NPU] Adapt qwen3-next W8A8 on NPU (#16164) shengzhaotian 2026-01-03 19:41:20 +08:00
  • 2c09de343e [diffusion] improve: skip loading vision module for text encoders (#16304) Mick 2026-01-03 19:30:45 +08:00
  • 38d48de93d [diffusion] CI: simplify warmup (#16303) Mick 2026-01-03 18:09:38 +08:00
  • 7f2fa2167b Tiny add --log-requests-target (#16338) fzyzcjy 2026-01-03 17:28:27 +08:00
  • d0fb24ee7b [Diffusion] Flux2 tp support (#16219) Xiaoyu Zhang 2026-01-03 17:02:27 +08:00
  • 65b0b5b248 fix(logging): use Display format for model_id instead of Debug (#16337) Simo Lin 2026-01-03 00:21:39 -08:00
  • 9a414b164c [Performance] Optimze the performance of Qwen25VL (#15640) Siyuan Chen 2026-01-03 15:15:36 +08:00
  • 5b4f790200 [diffusion] CI: add CI validation for diffusion model downloads (#16311) Alison Shao 2026-01-02 22:12:45 -08:00
  • d8ac5eecf7 [model-gateway] bug fix on module name (#16332) Chang Su 2026-01-02 21:56:18 -08:00
  • bdde949619 [HiCache] Add PP Support with suffix pp rank (#15175) Teng Ma 2026-01-03 13:49:21 +08:00
  • b23e7ed13c refactor(gateway): shorten logging targets from sgl_model_gateway to smg (#16328) Chang Su 2026-01-02 21:12:33 -08:00
  • 9821fae5c6 Tiny add explanations to realtime token metric labels (#16009) fzyzcjy 2026-01-03 12:36:55 +08:00
  • 078d96213a [FEAT] optimize tensor zmq transfer for multimodal inputs (#13592) siyu 2026-01-03 12:23:23 +08:00
  • 8b111b20c3 feature: improvements to CI failure monitor (#16272) Douglas Yang 2026-01-02 20:09:41 -08:00
  • 74a166cb86 [Fix] Only add SM90 and SM100 to check for auto-enabling TRT Allreduce Fusion (#16283) b8zhong 2026-01-02 19:43:17 -08:00
  • 888e126ac9 [diffusion] comment: fix typo (#16257) triple-mu 2026-01-03 12:42:04 +09:00
  • bb23a8fe77 [Tiny]Remove progress bar for fp8 ue8m0 quant when unneeded (#16177) Baizhou Zhang 2026-01-03 10:54:12 +08:00
  • 8b869e326c [AMD] feat: add DLLM support for AMD GPUs with LLaDA2 testing (#15560) sunxxuns 2026-01-02 18:41:11 -08:00
  • 6256936d09 Fix: Allow build-wheels to run when target_stage is set (#16324) Alison Shao 2026-01-02 17:58:11 -08:00
  • 62f73a8c81 [Auto Sync] Update engine.py (20260102) (#16317) Lianmin Zheng 2026-01-02 17:29:40 -08:00
  • 7c1b4b1c4c Support swa allocator page size > 1 (#16296) Ke Bao 2026-01-03 08:49:47 +08:00
  • 31ed68e7c1 [model-gateway] address embedding similarity threshold in ut (#16321) Simo Lin 2026-01-02 16:46:55 -08:00