Commit Graph

  • 8cdb7e1fd4 [CI] Add GPT-OSS test for SM120 (#20056) Mohammad Miadh Angkad 2026-03-07 03:43:04 +08:00
  • 759700c808 Fix SM120 triton_kernels MXFP4 block_k for GPT-OSS (#20040) Mohammad Miadh Angkad 2026-03-07 02:53:08 +08:00
  • de1a0afcbc [MUSA][10/N] Add GGUF support (#18357) R0CKSTAR 2026-03-07 02:50:35 +08:00
  • e8f2b80340 [diffusion] improve: improve code readability of DenoisingStage (#20003) JohnHerry 2026-03-06 23:23:44 +08:00
  • 54634b9a40 [Kernel] Dispatch exp/sin/cos through dtype_trait (#19798) xingsy97 2026-03-06 22:57:52 +08:00
  • 2d266c73ea Migrate renorm kernels from sgl-kernel to FlashInfer JIT (#18854) Johnsonms 2026-03-06 06:53:28 -08:00
  • 6d22c9f369 [Diffusion] Move hf kernels diffusion cuda kernels skills to SGLD (#20001) Xiaoyu Zhang 2026-03-06 22:16:06 +08:00
  • f7de9375ac [GDN][Qwen3-Next][Qwen3.5] Fuse fused_gdn_gating and fused_recurrent_gated_delta_rule_update in verify_target (#19775) Yuan Luo 2026-03-06 21:42:44 +08:00
  • e3b581ce6b [diffusion] fix: remove num_frames in wan2_1_t2v_1_3b_lora_1gpu test (#20009) Prozac614 2026-03-06 21:36:43 +08:00
  • 25e678d933 [diffusion] endpoint: add /server_info and /model_info endpoints for gateway discovery (#20020) Kangyan-Zhou 2026-03-06 05:36:13 -08:00
  • 550506894a [AMD] Upgrade aiter version (#19936) Thomas Wang 2026-03-06 18:31:04 +08:00
  • 84aaa69795 [AMD] Use bfloat16 for correction_bias in AITER FP8 path to avoid runtime dtype conversion for dsv3 (#19843) inkcherry 2026-03-06 16:57:12 +08:00
  • 27053aa5ed Fix MLA decode path returning unwritten (padded) rows (#19902) Clint 2026-03-06 02:54:29 -06:00
  • 0252ca8255 [Bugfix] Fix the bug blocking the startup of Llama-3.2-11b-Vision-Instruct (#19638) xdtbynd 2026-03-06 16:21:50 +08:00
  • 23eb24ddfe [AMD] Update/fix AMD CI workflow dispatch mechanism (#20014) YC Tseng 2026-03-06 10:27:44 +03:00
  • da27d9bff6 [Bug-Fix][EPD]: skip log waiting-image-req for zmq_to_tokenzer/mooncake (#19555) Zheng Wengang 2026-03-06 14:39:22 +08:00
  • 04e364d538 [V32] Enhance deepseek v32 related tests (#19985) Baizhou Zhang 2026-03-05 20:12:49 -08:00
  • be9a9e4819 refactor(multimodal/test): centralize model names and shared utilities in test_utils (#19354) Mook 2026-03-05 20:09:42 -08:00
  • 51e5dc845a Revert "[Kernel Slimming] Migrate NVFP4 kernels to JIT" (#20005) Baizhou Zhang 2026-03-05 19:40:00 -08:00
  • 6e5a2de354 [diffusion] fix: fix reading multiple prompts from prompt file (#19075) sushil Dubey 2026-03-06 08:53:31 +05:30
  • 9502369488 fix(grpc): add server-side keepalive options to prevent GOAWAY (#19986) Simo Lin 2026-03-05 18:56:35 -08:00
  • 5471e4a492 [NPU][Feature] eliminate dsv3 redundant rotary embed calculation (#19842) liupeng374 2026-03-06 09:02:14 +08:00
  • b912d7ae19 [OPT]Skip the first delayer to maximize the BS of the decoding. (#19836) chenxu214 2026-03-06 08:53:19 +08:00
  • 261be85ecc Support mrope_position_delta cache shadowxz109 2026-03-06 08:50:53 +08:00
  • 9ebffef1ef [FIX] NSA backend page_table overflow in speculative decoding target_verify (#19016) Xinyuan Tong 2026-03-06 00:04:58 +00:00
  • 13af7cbb02 fix: use consistent time denominator for throughput metrics in bench_one_batch_server (#19223) Ajay Anubolu 2026-03-05 15:58:17 -08:00
  • dd2bbe6d62 fix(grpc): use context.abort() with proper status codes instead of in-band errors (#19972) Chang Su 2026-03-05 14:53:18 -08:00
  • 46dced64ea Adjust padding size to improve triton_kernels moe performance (#19174) Qiaolin Yu 2026-03-05 14:50:40 -08:00
  • 346a4131cf [Spec] Refactor NaN/OOB checks to async maybe_detect_* with env-var control (#19899) kpham-sgl 2026-03-05 13:51:05 -08:00
  • 1c1712d8e5 [CI] Skip flashinfer-cubin reinstall when version matches (#19470) Alison Shao 2026-03-05 13:30:44 -08:00
  • b3cfad0a80 Add Ray actor support for scheduler process management (DP=1) (#17684) Xinyu Zhang 2026-03-05 13:21:23 -08:00
  • 07e7603c0c Update sgl-attn to include SWA decode optimizations (#19655) Minglei Zhu 2026-03-05 12:45:25 -08:00
  • ebb66cc1de [misc] Priority scheduling metrics cleanup (#19927) sglang-bot 2026-03-05 12:42:42 -08:00
  • ff6048fb9c rename nemotron reasoning parser (#19865) danielafrimi 2026-03-05 21:27:07 +02:00
  • e58391dd7d Add --json-log flag to enable structured JSON logging (#19968) Jonathan Lee 2026-03-05 14:24:12 -05:00
  • 41fd53fe37 Fix profile_activities parameter name in bench_one_batch_server_internal.py (#19954) Mohammad Miadh Angkad 2026-03-06 02:34:06 +08:00
  • d605d811fb update CODEOWNERS (#19969) Mick 2026-03-06 02:19:05 +08:00
  • 203cd8eb02 [AMD] [Z-Image-Turbo Day 0] Add Z-Image-Turbo nightly test for AMD GPUs (#19733) Michael 2026-03-05 08:36:17 -08:00
  • 73d272bddb Revised fix for HybridAttnBackend forward for linear attn (#19369) akhilg-nv 2026-03-05 08:05:35 -08:00
  • b5edab57f2 [AMD] CI - Add MI35x nightly/PR tests for kv-cache-fp8 and allreduce-fusion (DeepSeek) (#19834) YC Tseng 2026-03-05 18:09:57 +03:00
  • 0de0d74195 [EPD][Feat]support adaptive forward (#18118) Zheng Wengang 2026-03-05 21:12:30 +08:00
  • 806d41ab65 [quant] fix fp32 downcasting (#19844) StonyPort 2026-03-05 17:54:59 +08:00
  • 472eef4071 fa4 cleanup (#19727) Rain Jiang 2026-03-05 01:54:25 -08:00
  • c36de62bfc [diffusion] fix images/edit with 2 images (#17520) Chi McIsaac 2026-03-05 03:56:39 -05:00
  • dbc896f204 [Test] Enhance JIT kvcache store kernel test coverage (#19630) xingsy97 2026-03-05 16:17:15 +08:00
  • 727face6c2 [DLLM] Add initial radix cache support (#18724) Tiwei Bie 2026-03-05 15:24:09 +08:00
  • c1df359b44 Add XPU profiler activity support in benchmark code (#12981) Kalyan Kumar 2026-03-05 12:52:56 +05:30
  • 2bdd89a6cd [Kernel Slimming] Migrate NVFP4 kernels to JIT (#19437) Mohammad Miadh Angkad 2026-03-05 15:22:28 +08:00
  • 1bbfed0539 [misc] add env for http keep alive timeout (#19847) Yilong Zhao 2026-03-04 22:00:51 -08:00
  • 86c5617787 [BUG]: fix prevent illegal memory access in Mamba SSM tracking during EAGLE speculative verification (#19415) Chenxi Li 2026-03-04 21:13:21 -08:00
  • 10c65df48a [Bug] Fix lora tp bug on H200 (#19769) Baizhou Zhang 2026-03-04 20:11:02 -08:00
  • feda2b11c4 [AMD] Add AWQ AMD CI coverage and quantization platform compatibility docs (#19550) Bruce Changlong Xu 2026-03-04 19:50:55 -08:00
  • 0e6a64712a [bugfix] Fix PPMissingLayer AttributeError when Using PP (#19804) Xinyi Song 2026-03-04 19:48:15 -08:00
  • 198381d9ce Add SSL/TLS support for HTTP and gRPC servers (#18973) Kangyan-Zhou 2026-03-04 19:27:16 -08:00
  • 9c11a7ae40 [diffusion] fix: fix the frame interpolation testcase in CI regarding number of frames (#19659) Junhao Liu 2026-03-04 19:21:53 -08:00
  • fc53307ce9 [diffusion] hardware: SiluAndMul/RMSNorm/LayerNorm MUSA implementations (custom ops, 12/N) (#18583) R0CKSTAR 2026-03-05 11:10:57 +08:00
  • 9795b4cd5b [Diffusion] Open t5 encoder parallel folding for wan2.2 and mova video (#18493) Xiaoyu Zhang 2026-03-05 10:18:00 +08:00
  • e555a6c171 [feat] Enhance lora_update_weight_from_tensor for RL training (#19314) Ethan (Yusheng) Su 2026-03-04 18:10:42 -08:00
  • d8427d0156 [NPU][CI] Cache pytorch dependency in ci (#19754) monkeyLoveding 2026-03-05 09:24:37 +08:00
  • 0eb64c1e72 [smg] Extract tokenizer_path from /model_info into discovered labels (#19905) Kangyan-Zhou 2026-03-04 17:12:20 -08:00
  • 861d78635f [CI] remove itl testing due to unstable networking (#19904) Liangsheng Yin 2026-03-04 16:32:17 -08:00
  • 43bdee703e Fix Fp8 MTP layer a2a backend without EP. (#18515) Shu Wang 2026-03-04 18:28:10 -06:00
  • 33c92732f4 [Triton] Use dynamic loop bound in alloc_extend_kernel (#19898) Liangsheng Yin 2026-03-04 16:15:58 -08:00
  • a710b7d791 [Sarvam] Add inference support for Sarvam MoE LLMs (#18938) rakesh 2026-03-05 04:58:00 +05:30
  • 376dfb03f7 Fix issue 19717 by making qo_indptr uniform strided instead of packed (#19807) kpham-sgl 2026-03-04 15:27:10 -08:00
  • 28c931e1a5 feat: Priority-based scheduling optimization (including default priority, preemption toggle, priority-based metrics, etc.) (#17026) zhuxinjie-nz 2026-03-05 06:52:08 +08:00
  • 9457c049e1 [Qwen3.5] Enable MTP spec_v2 and add test for nvidia/Qwen3.5-397B-A17B-NVFP4 (#19391) hlu1 2026-03-04 14:01:25 -08:00
  • 0ee9d3c8e9 fix(grpc): send last chunk before completion during streaming (#19895) Chang Su 2026-03-04 13:21:21 -08:00
  • 329817e262 [AMD] Move get_global_server_args import out of CUDA-only block to fix NameError on AMD (#19866) Bingxu Chen 2026-03-05 02:23:42 +08:00
  • 1b76eb9361 [Doc] Update version references and add automation (#18409) Mohammad Miadh Angkad 2026-03-05 01:51:46 +08:00
  • 44208d2adf [vlm][minicpm] support input formats of processor output and embedding (#19614) Ken J 2026-03-04 09:11:12 -08:00
  • c03deb8175 Fix disagg PD bootstrap and KV transfer metrics (#19009) Kangyan-Zhou 2026-03-04 09:08:10 -08:00
  • 34c19a32c1 fix flaky test for test_kda_kernels (#19864) strgrb 2026-03-04 22:47:29 +08:00
  • 738ebfd330 KDA: fuse qkv conv and support stride for fused_sigmoid_gating_delta_rule_update_kernel (#19506) strgrb 2026-03-04 22:45:53 +08:00
  • 6910c1b281 [Feature][NPU]: add runtime support for GPTQ-quantized MoE models (#16364) YeChang Guo 2026-03-04 21:02:19 +08:00
  • c2b66d320d [HiCache] Add an env var to control transfer engine reuse (#19867) Shangming Cai 2026-03-04 20:36:32 +08:00
  • e33e833d11 update model names (#19870) amote-i 2026-03-04 19:27:37 +08:00
  • 88cfa6c11d [NPU]Releasing redundant memory of w13_weight and nz when the ascend_fuseep feature is enabled (#19813) chenxu214 2026-03-04 19:26:29 +08:00
  • 17119a697d Optimization: Reduce the number of D2H operations (#19424) sky 2026-03-04 16:32:42 +08:00
  • f07d668ba1 Remove flashinfer version argument from cu13 docker release workflow (#19862) Baizhou Zhang 2026-03-04 00:26:06 -08:00
  • 52dcade4aa Fix flashinfer bump workflow (#19855) Baizhou Zhang 2026-03-04 00:09:58 -08:00
  • 78ddf05afd [Fix] Install tomli in flashinfer bumping workflow (#19841) Baizhou Zhang 2026-03-03 23:12:22 -08:00
  • 09fa012ba7 Fix /health regression from early prebound socket listen (#19805) Mohammad Miadh Angkad 2026-03-04 15:00:46 +08:00
  • 115f879958 Helios: Real Real-Time Long Video Generation Model (#19782) Yuhao Yang 2026-03-04 14:58:04 +08:00
  • c287d9b645 chore: add flashinfer version bump workflow (#19837) Baizhou Zhang 2026-03-03 22:53:34 -08:00
  • 562c3ff2d0 [Feature] implement the standard multi-layer MTP for step3p5 (#18564) qwe 2026-03-04 14:48:53 +08:00
  • e9b5706545 [diffusion] feat: support torch compile for diffusers backend (#19673) DefTruth 2026-03-04 14:08:45 +08:00
  • c6850ac30c [AMD] Fix Qwen3-Coder-Next: Add missing k_scale/v_scale args to extend_attention_fwd in aiter_backend (#19736) Michael 2026-03-03 22:01:08 -08:00
  • 5972f97f11 Remove naive rotary forward overriding. (#19263) Jue Wang 2026-03-04 00:50:40 -05:00
  • ac1f07487a Fix triton alloc extend kernel (#19780) ybyang 2026-03-04 13:01:16 +08:00
  • 73bf2c5bdc [sgl]add pin_mem to remove cpu->gpu copy sync point (#19795) Bi Xue 2026-03-03 21:00:51 -08:00
  • b7f7df7ee6 [NSA] Fix line-too-long lint in can_nsa_prefill_cp_round_robin_split (#19829) sglang-bot 2026-03-03 20:34:22 -08:00
  • ca44aa25af Fix dp_attention crash when dp_size < tp_size in warmup dummy run (#19760) Yuhao Yang 2026-03-04 11:43:13 +08:00
  • da9dcbc906 [diffusion] fix: fix corrupted image editing outputs in Multi-GPU SP mode for FLUX.2-klein models (#19454) Ruihang Li 2026-03-04 11:35:46 +08:00
  • 6851613b93 [Bugfix] For cp: Fixed hang problem in prefix cache and kvcache support fp8 in-seq-split mode (#19656) Baidu-AIAK 2026-03-04 11:19:46 +08:00
  • 88290690f0 Add kpham-sgl into CI Permission list (#19819) kpham-sgl 2026-03-03 18:55:10 -08:00
  • 4348976f80 [Diffusion] Refactor diffusion benchmark/profile skill to reuse diffusion-perf skill and clarify profiling trigger (#19783) Xiaoyu Zhang 2026-03-04 10:54:42 +08:00
  • 82e7139c06 [VLM] Support cos sin cache for Ernie4.5-VL (#19743) Yuan Luo 2026-03-04 10:54:23 +08:00
  • 525d046990 [AMD] CI - new runner label for MI325 8gpu (#19815) YC Tseng 2026-03-04 05:45:23 +03:00
  • 115e9a1acd [Diffusion] Delete useless _ulysses_input_split func (#19786) Xiaoyu Zhang 2026-03-04 10:45:11 +08:00