Commit Graph

  • c1d1337afc [diffusion][Wan]: fix sparse attention backends being applied to cross-attention (#17596) Ratish P 2026-02-16 19:27:58 +05:30
  • b86c6491fa [Perf] ~9.5x faster Blackwell MXFP4 MoE weight loading (#18858) Mohammad Miadh Angkad 2026-02-16 19:47:09 +08:00
  • 4f0409f8aa [Model] Add Qwen3ForRewardModel and fix Qwen3ForSequenceClassification (#17992) Shivam jindal 2026-02-16 17:14:41 +05:30
  • de833f9e8e Revert "[diffusion]: Improve layerwise offload buffer reuse and shared-storage handling" (#18866) Mick 2026-02-16 18:00:58 +08:00
  • d0c94e136a [diffusion] logging: improve peak vram logging (#18865) Mick 2026-02-16 16:44:37 +08:00
  • ed22720c07 [JIT kernel] hd=512,1024 in JIT QK norm (cta based) (#17515) Yi Zhong 2026-02-16 03:07:24 -05:00
  • 206accd15d Fix GLM-4V processor registration when glm_ocr is unavailable (#18885) Alison Shao 2026-02-16 00:02:31 -08:00
  • 61da34ad0b [diffusion] fix: fix LoRA weight snapshot aliasing in unmerge logic (#18883) Changyi Yang 2026-02-16 02:39:45 -05:00
  • d85884ca57 Update ascend_npu_qwen3_5_examples.md (#18888) RAY 2026-02-16 15:24:27 +08:00
  • 86c181e335 Fix test_lora_qwen3 nightly failure: replace adapter with added_tokens (#18884) Alison Shao 2026-02-15 22:35:06 -08:00
  • f1efb46bdd fix: adding performance logging for nightly diffusion (#18023) Douglas Yang 2026-02-15 22:09:00 -08:00
  • 2050875424 fix: unifying docker image build pipeline (#18814) Douglas Yang 2026-02-15 22:05:55 -08:00
  • f554b3c27b Support dumping gradients, parameters, lazy values (#18881) fzyzcjy 2026-02-16 13:34:06 +08:00
  • 9a7d8d5eb0 Collect upper level metadata to dump output (#18880) fzyzcjy 2026-02-16 13:31:19 +08:00
  • 949792d0c6 Change dump output format to dict with value and metadata (#18879) fzyzcjy 2026-02-16 13:30:47 +08:00
  • 02816abc0d Flip dumper to disable by default and refactor environment handling (#18878) fzyzcjy 2026-02-16 13:29:32 +08:00
  • 5ddc84e33e [AMD] MORI-EP inter kernel type switch (#18437) Duyi-Wang 2026-02-16 12:59:39 +08:00
  • bc79a64d3a [Diff]: support SGLANG_TORCH_PROFILER_DIR environment variable for profiler log directory (#18454) Johnsonms 2026-02-15 20:47:29 -08:00
  • 0af9dcc407 [diffusion] refactor: refactor server_args adjust and validate logics (#18863) Mick 2026-02-16 11:49:06 +08:00
  • 78b4c9e248 [diffusion] fix: avoid saving output for warmup requests (#18867) Mick 2026-02-16 11:48:28 +08:00
  • 8a82c70297 [VLM] Optimize Ernie4.5-VL rotary embedding with fused triton kernel (#18856) Yuan Luo 2026-02-16 11:19:44 +08:00
  • 45715af50c fix: nightly whl dev date suffix (#18873) Douglas Yang 2026-02-15 18:57:37 -08:00
  • 0ffd0a3995 Nsa trtllm mla sparse fp8 support with Deepseek v3.2 NVFP4 (#18389) Rain Jiang 2026-02-15 17:29:54 -08:00
  • 8290171f52 [CI] Remove --mem-fraction-static 0.93 from gpt-oss test (#18869) Mohammad Miadh Angkad 2026-02-16 09:24:11 +08:00
  • d3bae71e3f Add claude skills for sgl-kernel and jit-kernel (#18855) Xiaoyu Zhang 2026-02-16 07:14:03 +08:00
  • fd5a45d5cf Update ascend_npu_support.rst (#18868) chenxu214 2026-02-16 01:41:38 +08:00
  • f2d72866e9 Create ascend_npu_qwen3_5_examples.md (#18864) chenxu214 2026-02-16 01:15:20 +08:00
  • 0d30896015 fix(sgl-kernel): use >= 120 for SM12x CUDA kernel dispatch (#18750) blake-snc 2026-02-15 08:44:47 -08:00
  • 5fc328465a fix(sgl-kernel): support CUDA 13 runtime preloading for DGX Spark (#18747) blake-snc 2026-02-15 08:43:04 -08:00
  • b79808bee2 Fix libnuma.so does not exsit (#15355) Mike Qiu 2026-02-16 00:37:50 +08:00
  • 48eac1b62d Improve profiler options for bench_serving (#16991) akhilg-nv 2026-02-15 08:36:01 -08:00
  • f759960ac2 feat: expose consistent_hashing policy in Python router CLI args (#17972) Blake Ledden 2026-02-15 08:33:24 -08:00
  • 597d17dd18 Use ephemeral nccl port via get_free_port() (#18009) Chanh Nguyen 2026-02-15 08:32:47 -08:00
  • 7a607c4900 fix_get_quant_method_in_fused_moe_condition (#18459) tjp_zju 2026-02-16 00:31:42 +08:00
  • 536ed3143b test: add test for Modelopt FP8 on SM90 (#18463) Zack Yu 2026-02-15 08:29:37 -08:00
  • b2f74d660a fix: add SM110 (Jetson AGX Thor) to Blackwell capability check (#18787) WiwilZ 2026-02-16 00:26:58 +08:00
  • 57f7e06cb9 fix: update Blackwell log/error messages to include SM12x (#18751) blake-snc 2026-02-15 08:23:51 -08:00
  • 07a24f1a38 update pre-commit config (#18860) SoluMilken 2026-02-16 00:18:31 +08:00
  • f7603203b0 Enable DeepGemm fast warmup in CI to prevent cold-cache timeouts (#18823) Alison Shao 2026-02-15 08:02:30 -08:00
  • 4ef8ece08a feature: adding build commit to sgl kernel workflow (#18853) Douglas Yang 2026-02-15 07:28:43 -08:00
  • ddfe147377 [diffusion]: Improve layerwise offload buffer reuse and shared-storage handling (#18611) Ratish P 2026-02-15 19:47:51 +05:30
  • 4e162d4b1b change npu.dockerfile (#18835) chenxu214 2026-02-15 20:43:15 +08:00
  • 3feb48139e [diffusion] quant: add support for svdquant and nunchaku (#18549) Mick 2026-02-15 20:43:00 +08:00
  • 88010e9601 [AMD] Fix nightly 1-GPU test failures and bench_serving regression (#18761) Michael 2026-02-15 04:36:47 -08:00
  • 90555a0228 Add missing dumper tests (#18859) fzyzcjy 2026-02-15 18:42:57 +08:00
  • 4c7f986c6b Extract dumper and prefill delayer tests common utils (#18857) fzyzcjy 2026-02-15 18:33:23 +08:00
  • b992828ad2 fix: fix bug on kimi2.5 with dp2 and tp4 (#18604) haowen-han 2026-02-15 16:32:13 +08:00
  • 274bf6607a [diffusion] fix: enable torch.compile for UlyssesAttention (#18840) Ratish P 2026-02-15 13:24:27 +05:30
  • ad1bdb93df perf: add minimax-2.5 fused_moe tuning config for h20 (#18833) zhangxiaolei123456 2026-02-15 15:46:56 +08:00
  • 922fbc21e2 [Perf] Tune MiniMax M2 fused moe kernel on H100 GPU (#18851) jackey hua 2026-02-15 02:30:52 -05:00
  • 944a9f6fcf Fix/qwen3 5 amd rope cutedsl fallback (#18753) andyluo7 2026-02-14 22:09:44 -08:00
  • 91230dcca8 [FIX] Correct JIT kernel compilation on newer GPUs with outdated driver metadata. (#18496) muse-coder 2026-02-15 12:14:39 +08:00
  • 1ce3420784 Model: Support IBM Granite (Dense/Mamba + MoE) (#18040) Bhavneek Singh 2026-02-15 12:24:41 +09:00
  • 4cf4f0859f [Doc] Convert the speculative decoding notebook to markdow (#18395) shuwenn 2026-02-15 10:18:56 +08:00
  • b33769786f [Auto Sync] Update grpc_request_manager.py, tokenizer_manag... (20260214) (#18838) Lianmin Zheng 2026-02-14 18:12:32 -08:00
  • 190fa8246f Fix model loading for DeepSeek-V3.2-AWQ (#16907) Guangda Liu 2026-02-15 09:39:53 +08:00
  • 8b2020584c [Auto Sync] Update test_deterministic.py (20260214) (#18839) Lianmin Zheng 2026-02-14 17:19:30 -08:00
  • b1b69ae0a9 Add CI permissions (#18847) Mohammad Miadh Angkad 2026-02-15 08:24:36 +08:00
  • 4067d9487d [diffusion] feat: opt vae decode with channels_last_3d (#18540) Xiaoyu Zhang 2026-02-14 23:19:45 +08:00
  • c29394e3c8 [kernel slimming] Move fast_hadamard_transform to jit_kernel (#18475) Xiaoyu Zhang 2026-02-14 23:06:21 +08:00
  • ae95869292 Enable SGLANG_ENABLE_SPEC_V2 for nightly speculative decoding tests (#18719) Kangyan-Zhou 2026-02-14 23:00:33 +08:00
  • 92cdd398cd feat: Support mrope_section with rope_type: "yarn" (#13313) Raayan Dhar 2026-02-14 06:51:44 -08:00
  • f51e9d9ca1 Add ci test for ring model (#18829) Ke Bao 2026-02-14 22:20:23 +08:00
  • c8aa2a6534 Fix dsv32 encode_messages (#18126) ybyang 2026-02-14 16:44:13 +08:00
  • 34132d6da5 Kernel: optimize decoding metadata in NSA multi-spec backend with fused kernels (#17554) Johnsonms 2026-02-14 00:40:15 -08:00
  • 38473f8ee0 [AMD] Fix sgl-model-gateway Build Errors in ROCm Docker Release (#18836) Bingxu Chen 2026-02-14 16:07:26 +08:00
  • e2eb5bf28d [DLLM] Update CODEOWNERS for diffusion LLM (#18834) Zehuan Li 2026-02-14 15:31:42 +08:00
  • fa0ef6e4f7 [VLM][LLM] Optimize fused_moe triton kernel tma (#18782) Yuan Luo 2026-02-14 14:35:26 +08:00
  • f6c18c3a85 Fix/partial gen from waiting queue miss metadata (#17610) JD 2026-02-13 19:04:08 -08:00
  • 45a4697d45 [diffusion][MUSA] fix: MUSA platform breakage caused by PR #13662 (#18456) R0CKSTAR 2026-02-14 11:00:39 +08:00
  • 8ef3e3d56b Fix CI concurrency collision between scheduled runs and fork PRs (#18826) Alison Shao 2026-02-13 18:48:31 -08:00
  • 066b0b70d9 Handle abort for retracted requests in disagg decode prealloc queue (#18705) qmzznbxhl 2026-02-14 10:39:39 +08:00
  • bd39de7d5e [Env] centralize hicache vars in environ.py (#17204) shuwenn 2026-02-14 10:02:31 +08:00
  • dcea74d63f Add timeout abort kits for normal / eagle. (#18815) Liangsheng Yin 2026-02-13 17:57:30 -08:00
  • 4474fb98b4 [PD-Disagg] Fix double free when prebuilt batch is aborted. (#18822) Liangsheng Yin 2026-02-13 17:46:35 -08:00
  • ab0fb248fd feat: add SGLANG_DISTRIBUTED_INIT_METHOD_OVERRIDE env var (#18743) Leon Gao 2026-02-13 17:37:33 -08:00
  • 8be18c655d [Perf] refactor piecewise cuda graph support of Qwen3-Next (#17613) Minglei Zhu 2026-02-13 17:30:50 -08:00
  • 3a1c388b43 Update performance dashboard for nightly tests (#18824) Kangyan-Zhou 2026-02-14 09:28:28 +08:00
  • 3299c4f9c1 [CI] feat: add early exit to wait_for_server when process dies (#18602) shuwenn 2026-02-14 08:46:09 +08:00
  • eccf875d49 [CI] Revive 8-GPU trace upload in nightly test workflow (#18820) Kangyan-Zhou 2026-02-14 08:37:08 +08:00
  • 1be41e9036 [FlashInfer] Bump FlashInfer version from 0.6.2 to 0.6.3 (#18448) Mohammad Miadh Angkad 2026-02-14 07:43:33 +08:00
  • 710d873ba6 Update notified user in post_ci_failures_to_slack.py (#18817) Kangyan-Zhou 2026-02-14 06:48:56 +08:00
  • 191d354f53 fix double-free kv cache for requests that have already finished and been freed during preemption (#18694) JD 2026-02-13 13:17:44 -08:00
  • 008ea46af1 [Auto Sync] Update loader.py, weight_utils.py (20260213) (#18779) Lianmin Zheng 2026-02-13 12:22:50 -08:00
  • 4c6afbeeaa [bugfix] fix mamba slot leak when scheduling fails with radix cache (#15840) (#16067) Qi Jia 2026-02-13 23:43:57 +08:00
  • 8b4c364960 refactor context parallel state (#17213) dongjiyingdjy 2026-02-13 23:18:17 +08:00
  • 0012d6a4eb [Kernel Slimming] Migrate GPTQ-Marlin repack kernel to JIT (#18543) Linyu Wu 2026-02-13 22:29:22 +08:00
  • 37273408eb [diffusion] chore: use batched P2P ops in VAE parallel decoding (#18728) Mick 2026-02-13 22:11:20 +08:00
  • acc940d302 [diffusion] fix typo (#18790) triple-mu 2026-02-13 21:59:39 +08:00
  • 9a32f8ccb9 [CI] Move test_load_lora_from_tensor test to H100 (#18797) Baizhou Zhang 2026-02-13 21:28:00 +08:00
  • 07633349c9 [diffusion] fix: webui task_type check (#18462) R0CKSTAR 2026-02-13 21:19:16 +08:00
  • efdd676d56 [diffusion] refactor: merge redundant default_dtype and param_dtype parameters in FSDP loader (#18789) Mick 2026-02-13 21:18:02 +08:00
  • 98ad284ebf Added cuda availability guard (#18480) Kaixi 2026-02-13 13:18:34 +01:00
  • a0ebaa6498 Cleanup debug log for Ring model (#18793) Ke Bao 2026-02-13 18:36:20 +08:00
  • a9d59776cc Enhence gsm8k test (#18791) Ke Bao 2026-02-13 18:08:57 +08:00
  • eacab2868a Adjust mamba cache allocation (#18786) Ke Bao 2026-02-13 18:06:23 +08:00
  • a6c4b52ac5 Cleanup unused rerun stages (#18788) Ke Bao 2026-02-13 17:44:42 +08:00
  • e4b2b57620 [schedule] Fix streaming return of customized_info (#18654) Yinghai Lu 2026-02-13 01:19:16 -08:00
  • 356e338607 [diffusion] feat: support SparseVideoGen2 attention backend (#17507) Xinwei Qiang 2026-02-13 16:20:46 +08:00
  • d97eb111a3 Support LingV2_5 model (#18598) ant-yy 2026-02-13 16:09:15 +08:00