Commit Graph

  • 51b3ed02ca Fix bug in symm mem pre-allocation default (#19082) Nicolas Castet 2026-02-20 21:51:28 -06:00
  • afd91e8782 [DSv32] Fix MTP and CP compatability (#19062) Vladislav Nosivskoy 2026-02-21 06:10:34 +03:00
  • 463baafe10 [Auto Sync] Update batch_invariant_ops.py (20260221) (#19098) Lianmin Zheng 2026-02-20 18:27:31 -08:00
  • 2928dfb8fa [Auto Sync] Update bench_one_batch_server_internal.py (20260221) (#19097) Lianmin Zheng 2026-02-20 18:19:14 -08:00
  • 84c67c8be0 Refactor graph input buffers (#18991) Cheng Wan 2026-02-20 18:09:31 -08:00
  • b2573fe426 Upd: CODEOWNERS (#19055) HAI 2026-02-20 15:51:53 -08:00
  • 4bffd3a232 [GPT-OSS] support fp8 online quantization for gpt-oss bf16 (#18988) Minglei Zhu 2026-02-20 14:16:57 -08:00
  • 96bae2355e Add generated-shared-prefix dataset in bench_one_batch (#18986) Qiaolin Yu 2026-02-20 13:33:10 -08:00
  • ab18734375 [feat] feat: support swa in trtllm_mha (#18970) 0xNullPath 2026-02-21 01:39:29 +08:00
  • fbb6098487 [AMD] support two batch overlapping for mori ep (#17953) billishyahao 2026-02-21 00:45:55 +08:00
  • 38ee749dd9 Fix adjust_num_token_non_padded_for_attn_tp returning CPU tensor (#19051) Cheng Wan 2026-02-20 07:23:38 -08:00
  • 3358ba8945 [Fix] Run FlashInfer autotune on non-default stream for NCCL 2.29+ compatibility (#18987) Nicolas Castet 2026-02-20 09:21:38 -06:00
  • 52852404c8 [Fix] DO NOT skip save_kv_cache for dllm (#19020) DarkSharpness 2026-02-20 23:20:29 +08:00
  • f23a23cc05 Fix NSA FP8 KV cache path for both-trtllm MHA one-shot (#18931) Mohammad Miadh Angkad 2026-02-20 22:00:09 +08:00
  • 8d789b5c3d [diffusion] feat: support nunchaku for Z-Image-Turbo and flux.1 (int4) (#18959) Mick 2026-02-20 21:16:08 +08:00
  • 7d953440ec [jit kernel] Support per_token_group_quant_8bit jit kernel (#18905) Yuan Luo 2026-02-20 21:01:05 +08:00
  • 38a69652e6 [diffusion] logging: log available mem when each stage starts in debug level (#18998) Mick 2026-02-20 19:57:06 +08:00
  • 0d20cf5a66 Fix lint on main (#19054) fzyzcjy 2026-02-20 15:45:24 +08:00
  • 77fdb6af81 feature: docker patch workflow (#19025) Douglas Yang 2026-02-19 23:37:40 -08:00
  • b59a22f781 fix lint on main (#19052) Cheng Wan 2026-02-19 23:30:57 -08:00
  • b0786cdf94 [AMD] Replace msgpack with msgspec in MORI-IO (#19007) Duyi-Wang 2026-02-20 15:04:15 +08:00
  • 8541b1118d [Fix][Qwen3.5] Pass max_mamba_cache_size to mamba pool in disaggregation decode path (#19002) YAMY 2026-02-19 22:31:26 -08:00
  • 295bc17576 Feature/sdar support (#19044) chengshuang18 2026-02-20 13:58:15 +08:00
  • 046ef0aa35 Support using SGLang port in dumper (#19038) fzyzcjy 2026-02-20 12:30:24 +08:00
  • 2fecc2c075 Support resetting and enhance HTTP endpoints for dumper (#19046) fzyzcjy 2026-02-20 12:29:09 +08:00
  • 503bf3047a Enhance configure and env parsing in dumper (#19034) fzyzcjy 2026-02-20 12:28:10 +08:00
  • df995aab56 Support filtering labels in dumper (#19018) fzyzcjy 2026-02-20 12:27:12 +08:00
  • 261bca3c58 Support captured dump output and console output control in dumper (#19017) fzyzcjy 2026-02-20 12:26:24 +08:00
  • fc1500adc6 Hint users when wrongly execute it with partial ranks in dumper (#19014) fzyzcjy 2026-02-20 12:25:54 +08:00
  • b41d412c3d Support cleanup previous dumps in dumper (#19013) fzyzcjy 2026-02-20 12:25:21 +08:00
  • 13a4a0406e Fix flashinfer autotune to only wrap run_once() (#19004) Cheng Wan 2026-02-19 20:02:21 -08:00
  • 64bca5315f Fix long prompt KV allocation by falling back to torch native APIs when exceeding Triton tensor limit (#18250) Cheng Wan 2026-02-19 19:15:05 -08:00
  • 99df920cdb Register tensors with symmetric memory for qwen (#18643) Nicolas Castet 2026-02-19 19:32:32 -06:00
  • 73a7f0d049 Revert "Add SDAR model support" (#19032) Cheng Wan 2026-02-19 16:03:56 -08:00
  • db34c1cbfb Tiny remove duplicate coredump env injection (#19023) Liangsheng Yin 2026-02-19 13:26:30 -08:00
  • 5ff5aa6923 [spec v2]Fix torch gc of future indices (#18958) Liangsheng Yin 2026-02-19 11:38:25 -08:00
  • 44ab752b7a Add SDAR model support (#18318) chengshuang18 2026-02-20 03:20:32 +08:00
  • 3207427d6d [diffusion] CI: enable warmup as default (#19010) Mick 2026-02-19 23:27:23 +08:00
  • d73f06f091 [diffusion] chore: improve memory usage on consumer-level GPU (#18997) Mick 2026-02-19 21:59:49 +08:00
  • 963def7f26 Move lora request validation to tokenizer_manager from server (#18962) satyamk7054 2026-02-19 05:03:19 -08:00
  • d07e8aa4a3 [Diffusion] [NPU] Enable profiler on NPU (#17807) Makcum888e 2026-02-19 15:33:51 +03:00
  • e21fc78dbd [diffusion] fix: fix rank used in parallel executor when enable_cfg_parallel is false (#18975) Prozac614 2026-02-19 20:12:24 +08:00
  • 19aa19b111 [diffusion] refactor: refactor diffusion triton kernels (#18966) Xiaoyu Zhang 2026-02-19 17:03:44 +08:00
  • 48642d5384 [RadixTree][4/N Refactor]: Move available_and_evictable_str to individual radix cache classes (#17852) pansicheng 2026-02-19 17:03:15 +08:00
  • 82a0bafc1c Feat/add fi selective state update kernel call (#18070) shaharmor98 2026-02-19 10:56:06 +02:00
  • 0be30d4b0d Fix PCG MoE Error (#17739) Yuwei An 2026-02-19 00:48:06 -08:00
  • bba2fc49a1 [Qwen3.5] Enable nvfp4 checkpoint (#18937) hlu1 2026-02-18 20:24:05 -08:00
  • 443b1a88d1 Add batched zero copy to NIXL backend (#18850) hxie 2026-02-18 16:31:02 -08:00
  • 462267982b [AMD] Fix mi35x dsv32 mtp nightly (#18978) Bingxu Chen 2026-02-19 08:23:17 +08:00
  • e2fccb2ee0 Fix flaky Qwen3-Next KL divergence tests by reverting mamba slot release (#18910) Alison Shao 2026-02-18 15:55:16 -08:00
  • 2f592c3b18 [Doc] Add flashinfer_deepgemm to --fp8-gemm-backend (#18982) Mohammad Miadh Angkad 2026-02-19 03:45:47 +08:00
  • 4f980f6f23 [Feature] Implement update_weights_from_disk for SGLang-D (Diffusion … (#18306) Mengyang Liu 2026-02-18 11:24:07 -08:00
  • 150ed881be [4/N] Quantization Refactor: Quark MoE schemes (#18252) Tamir Baydasov 2026-02-18 19:44:30 +03:00
  • 9c5aae4df5 [Fix] Add lora tied lm head support (for Qwen2.5, Gemma, etc model need) (#18634) Ethan (Yusheng) Su 2026-02-18 09:34:51 -07:00
  • 5a7ae059e3 Add DP ViT support for Kimi K2.5 (#18689) Yuhao Yang 2026-02-18 23:03:07 +08:00
  • eb6ff4a940 [diffusion] fix: refactor task resolution logic in benchmark function for multimodal generation (#18948) zijiexia 2026-02-18 04:57:42 -08:00
  • 0215d47007 [AMD] ROCm7.2: Add /sgl-workspace/aiter to PYTHONPATH (#18972) HAI 2026-02-18 02:21:39 -08:00
  • 90d5e27f79 Enable fa3 PDL by compiling it with corresponding flags (#18756) Qiaolin Yu 2026-02-18 01:12:05 -08:00
  • 390c154306 [Tiny fix] Super tiny fix mul_add naive forward bug (#18964) Xiaoyu Zhang 2026-02-18 16:18:43 +08:00
  • 513c12d23f Remove unused fast-hadamard-transform PyTorch extension sources (#18927) Xiaoyu Zhang 2026-02-18 15:51:07 +08:00
  • 934b36693c Reasoning models fix docs (#18963) HAI 2026-02-17 23:05:55 -08:00
  • 420a611275 [diffusion] refactor: unify SamplingParams construction and improve DiffGenerator return types (#18928) Mick 2026-02-18 14:56:58 +08:00
  • 9d138685c1 [Refactor] Fix test and clean up hicache code (#18555) DarkSharpness 2026-02-18 14:37:46 +08:00
  • 95c44cea29 [feat] Add return_routed_experts param to async_generate for parity with generate (#18508) William Arnold 2026-02-18 15:11:19 +09:00
  • ac0e493329 feat: add nsa and swa disagg support with nixl (#18939) Neal Vaidya 2026-02-17 21:13:26 -08:00
  • 2d85f01d43 Revert "Fix generated-shared-prefix bench_serving" (#18956) Liangsheng Yin 2026-02-17 20:43:55 -08:00
  • fa5698d791 feat: [Qwen3.5] Support block-wise FP8 quantization and model adaptation (#18926) Zheng Li 2026-02-18 11:44:25 +08:00
  • 83e24e2eb4 Expose priority parameter in Engine.generate() and Engine.async_generate() (#18944) Yan Ru Pei 2026-02-17 19:14:22 -08:00
  • 34d975b18f Fix eval tests not capturing server launch failures (#18886) Alison Shao 2026-02-17 15:59:03 -08:00
  • e02a9bec8d Refactor sampler: Use a better hash function for deterministic sampling and clear dispatch for probs/logprobs/logits sampling paths (#18915) Lianmin Zheng 2026-02-17 15:41:23 -08:00
  • 83a475e8d7 feat: add cuda core dump CI warpper (#18909) Liangsheng Yin 2026-02-17 14:49:26 -08:00
  • 9a7d6be567 cleanup prefill metrics logging to fix dp-attn metrics (#18778) Ratish P 2026-02-18 04:05:35 +05:30
  • 355127c2e9 Fix benchmark_sglang_fused_moe_triton.py (#18940) satyamk7054 2026-02-17 14:25:37 -08:00
  • 3c601db031 Fix generated-shared-prefix bench_serving (#18769) Qiaolin Yu 2026-02-17 14:00:22 -08:00
  • 48fcd62d1f fix(glm-image): single-GPU T5 config + SP support for 4D latents (#18… (#18739) Nickcp39 2026-02-17 16:04:32 -05:00
  • aeca7d348c [3/N] Quantization Refactor: ModelSlim MoE schemes (#17993) Tamir Baydasov 2026-02-17 21:38:27 +03:00
  • 10569d04bb [diffusion] update code owner (#18495) ronnie_zheng 2026-02-18 01:36:06 +08:00
  • bf08d3f43c [gRPC] Fix scheduler startup broken by context parallel refactor (#18933) Simo Lin 2026-02-17 08:52:11 -08:00
  • 504b2c58cf [diffusion] improve: improve torch.compile for MOVA (#18914) triple-mu 2026-02-18 00:47:38 +08:00
  • bf52388354 [PCG] support piecewise cuda graph for kimi-linear model (#18849) Minglei Zhu 2026-02-17 07:31:12 -08:00
  • bfe34c90ff Revert "[diffusion] operator: unify rotary embedding impl" (#18929) Mick 2026-02-17 22:56:04 +08:00
  • 2aa0db7d9c [Diffusion] [NPU] Fix CI run (#18921) Makcum888e 2026-02-17 16:54:19 +03:00
  • 8bb1037796 ROCm use rotary_embedding from sgl-kernel (#18920) HAI 2026-02-17 03:00:37 -08:00
  • 5f81ec1ad5 [Diffusion] Fix get model name when model local path end with "/" (#18918) Makcum888e 2026-02-17 13:19:54 +03:00
  • f6cc02489f [diffusion]: fix sparse video gen 2 backend being applied to cross-attention (#18900) Ratish P 2026-02-17 15:47:46 +05:30
  • b158f5d4a2 Revert "[AMD] Fix RotaryEmbedding crash on AMD/ROCm (regression from #17934)" (#18922) HAI 2026-02-17 01:07:50 -08:00
  • 14c95d255c [Diffusion] [NPU] [Doc] Add NPU documentation for sglang-diffusion (#18894) Makcum888e 2026-02-17 10:12:20 +03:00
  • 899e2be7d0 [TBO] fix cuda graph intermittently becomes disabled bug (#18320) billishyahao 2026-02-17 14:18:57 +08:00
  • 5e3103a787 [AMD] Fix RotaryEmbedding crash on AMD/ROCm (regression from #17934) (#18903) Michael 2026-02-16 20:59:40 -08:00
  • 90a0d66e1e [Tiny] Fix assert syntax warning in compressed_tensors_w4a4_mxint4_moe.py (#18899) Mohammad Miadh Angkad 2026-02-17 12:54:30 +08:00
  • 7e41ac6c8d Skip flaky test_tool_choice_required_non_streaming for Mistral (#18889) Alison Shao 2026-02-16 20:50:55 -08:00
  • d5307ce022 [misc] adding metadata field in UpdateWeightFromDiskReqInput (#18821) Yilong Zhao 2026-02-16 20:14:15 -08:00
  • 26b2c63d03 [diffusion] operator: unify rotary embedding impl (#18164) triple-mu 2026-02-17 12:02:48 +08:00
  • b21390f8f3 Adapt the Qwen2Model._update_causal_mask for transformers==4.57.1 (#18774) pansicheng 2026-02-17 10:20:41 +08:00
  • 50ca24aebb [diffusion]: fix scheduler crash on ZMQ messages with unexpected frame counts (#17890) Ratish P 2026-02-17 07:15:05 +05:30
  • f9c3def7fe Fix CI: add flashinfer --download-cubin to install dependencies (#18887) Alison Shao 2026-02-16 13:50:10 -08:00
  • 1b659bcb08 Fix GLM-5 fused shared expert (#18804) Frank Minors 2026-02-17 03:50:39 +08:00
  • 0ff24159a5 Fix modelopt FP8 create weights (#18447) danielafrimi 2026-02-16 18:59:50 +02:00
  • eba6af385d [2/N] Quantization Refactor: Compressed tensors MoE schemes (#17503) Tamir Baydasov 2026-02-16 18:03:51 +03:00
  • 1b3513a7e4 refactor FAKE transfer backend and remove --disaggregation-decode-enable-fake-auto parameter (#18345) Estrella-xx 2026-02-16 22:27:02 +08:00