Commit Graph

  • ee5ccde0ad support fused_moe_triton and moe_sum_all_reduce kernel fusion[reduce … (#19672) xieminghe1 2026-03-04 10:30:33 +08:00
  • d22c6a3847 fix: Properly return abort error for streaming requests if the abort is triggered by scheduler (#19357) Charles Chen 2026-03-03 17:18:15 -08:00
  • eb6bcc5c86 [CI] Register test_quant_config_parsing.py in CI suite (#19809) Alison Shao 2026-03-03 16:53:31 -08:00
  • e2af840c3d Various SM120 improvements (#19721) Brayden Zhong 2026-03-03 19:46:13 -05:00
  • a69b943356 [SGLang-Diffusion] Add offline throughput benchmark script for multi-modal models (#18154) Hao Jin 2026-03-03 16:39:46 -08:00
  • c18cff4f90 [CI] Add DeepGEMM warmup to stage-c-test-deepep-4-gpu (#19806) Alison Shao 2026-03-03 16:18:12 -08:00
  • 441045a7bf [AMD] Fix EAGLE3 speculative decoding with aiter attention backend (#19362) Hubert Lu 2026-03-03 16:12:13 -08:00
  • e6411ba315 Increase max_concurrent_jobs in job queue (#19797) Eric Zhang 2026-03-03 15:54:51 -08:00
  • c7ffbf25e9 [CI] Fix rerun-ut workflow: add DeepEP install, RDMA env, Blackwell detection (#19803) Liangsheng Yin 2026-03-03 15:17:16 -08:00
  • 753da27535 [Bugfix] fix parse_lscpu_topology bug (#18520) Jiayi Yan 2026-03-04 07:15:36 +08:00
  • 069c7e4188 Fix CI failures (#19303) Cao E 2026-03-04 07:04:44 +08:00
  • b8c71f895e Add tuned triton==3.5.1 h200 tp2, tp4 for qwen 3 next (#15948) Yi Zhong 2026-03-03 17:47:19 -05:00
  • 0c760c4cd7 Add tuned triton==3.5.1 b200 tp2, tp4 for qwen 3 next (#15917) Yi Zhong 2026-03-03 17:47:05 -05:00
  • fb37c0a400 [args] Add Expert Parallelism Argument To SRT Runner (#18492) Jonah Bernard 2026-03-03 17:16:35 -05:00
  • f7897def96 [Feature] Improve weight loading log (#18651) Praneth Paruchuri 2026-03-04 03:46:13 +05:30
  • 9305f0e58d Support triton_kernels for GPT-OSS on SM120 (#19718) Brayden Zhong 2026-03-03 17:14:01 -05:00
  • 1135e214b3 [CI] support /rerun-ut command in slash handler (#19800) Liangsheng Yin 2026-03-03 14:10:49 -08:00
  • 5b2e2750b5 Enable XQA for SM90 and SM120 (#17115) Sam (Kesen Li) 2026-03-04 06:09:44 +08:00
  • dc92f88a21 Enhance bench_multiturn.py with OpenAI API support and richer metrics (#19724) Kangyan-Zhou 2026-03-03 13:48:04 -08:00
  • f749802402 [Score API][18132] return token usage in Score API response (#18381) Guy Stone 2026-03-03 16:45:35 -05:00
  • b0f26698f5 feat(benchmark script): add similar to vllm --ready-check-timeout-sec parameter (#15466) almaslof 2026-03-03 23:44:38 +02:00
  • ac2819c81f Fix assertion tolerance for bf16 precision in triton attention UT (#17461) Rahul Vijayaraghavan 2026-03-04 03:13:58 +05:30
  • 85ab6a7f54 cli: Add lazy imports and fail-fast config validation (RFC #9853) (#19368) Karthik Koralla 2026-03-03 16:03:49 -05:00
  • d6ac5f23cc [Docs] Add GDN attention backends matrix documentation (#19755) zwang86 2026-03-03 13:00:34 -08:00
  • cedb86a950 Feature:Reserve HTTP server port before model loading to immediately detect port conflicts instead of failing after several minutes of model loading. (#17754) xrwang8 2026-03-04 03:59:18 +08:00
  • 2e1b9e2547 Fix routed_dp_rank boundary validation (#19762) doujiang24 2026-03-04 03:55:15 +08:00
  • 05f68e1230 [AMD] Fix the hipDeviceGetName issue in ROCm based docker images (#19440) Hubert Lu 2026-03-03 11:42:48 -08:00
  • 85f7a0aa30 feat: support Kimi K2.5 for Eagle3 (#19689) yefei12 2026-03-04 02:41:15 +08:00
  • 95e4a25b17 [fix typo]: funtion -> function (1 line change) (#19790) SoluMilken 2026-03-04 00:57:02 +08:00
  • daabfe7e4c [hotfix] fix apply function name compressed_tensors_w4a4_mxint4_moe (#19713) Tamir Baydasov 2026-03-03 16:50:37 +03:00
  • c6377bbbca feat(gdn): add FlashInfer K-last SSM layout support for GDN prefill and decode for Hopper (#18361) xutizhou 2026-03-03 20:30:48 +08:00
  • d939e26585 [model gateway][0/N] router EPD support: add encoder grpc server backend support (#16552) Jasonzhang517 2026-03-03 19:38:15 +08:00
  • facde4c6d3 [PD] Enable all CP ranks for KVCache transfer (#19765) Shangming Cai 2026-03-03 19:35:21 +08:00
  • 365ca1edb5 [NPU] bugs fix: fix a condition bug when using speculative inference on Qwen3 and Qwen3 moe (#19532) shengzhaotian 2026-03-03 17:59:25 +08:00
  • 666caaf9ce [Tool Call] Stream DeepSeek-V3.2 function call parameters in JSON format. (#16091) Muqi Li 2026-03-03 17:46:29 +08:00
  • 4c95953b77 Fix/nemotron mtp quantaized (#19433) Shaun Kotek 2026-03-03 11:07:46 +02:00
  • af0d35b224 Fix: Reject requests with a duplicate request ID which can cause server crash/hang (#19035) Charles Chen 2026-03-03 00:33:25 -08:00
  • 6af0448cc9 [Bugfix] Catch errors when DeepSeek-V3.2 generates malformed JSON (#18174) Muqi Li 2026-03-03 16:10:07 +08:00
  • 7a2d3df96f Apply default stream to priority 0 in scheduling. (#16438) Liangsheng Yin 2026-03-03 00:05:27 -08:00
  • 07b8d763ef feat: Add FP8 KV cache support for Triton attention backend (#18882) Zack Yu 2026-03-02 23:38:34 -08:00
  • 62480ebb1b [SGLang-Diffusion] Fix custom op fake impl missing eps default for torch.compile (#19725) 赵晨阳 2026-03-02 23:24:36 -08:00
  • 3c01b44700 [Fix] NPU deepep hccl buffer and fix IPC safe check (#17804) 1StepForever 2026-03-03 14:56:06 +08:00
  • 43a9249e7c CI: update CI_PERMISSIONS.json (#19744) Mick 2026-03-03 14:34:41 +08:00
  • dbf1247fe0 Add KimiK2Detector with tool interruption support (#19696) Xinyuan Tong 2026-03-03 06:04:49 +00:00
  • 63003a39cf [BUG] Support tuple hidden_states from fused MXFP4/FP8 quantization (#19643) Yuzhen Zhou 2026-03-02 20:39:06 -08:00
  • 0abb9f4176 Piecewise Cuda Graph Docs (#19738) Yuwei An 2026-03-02 19:51:17 -08:00
  • 6b8e62f94f [AMD] [Qwen 3.5 Day 0] Add Qwen 3.5 nightly accuracy tests (#19479) Michael 2026-03-02 19:42:42 -08:00
  • 060720c573 [AMD] AMD new CI runner (#19739) YC Tseng 2026-03-03 06:17:12 +03:00
  • fe9d85d93c Fix CompressedTensorsMxInt4MoE abstract method and relax GPQA baseline (#19726) Alison Shao 2026-03-02 19:03:21 -08:00
  • 8dfb6e1684 [HiCache] fix compatibility bugs with eagle and HiCacheStorage (#19570) huangtingwei 2026-03-03 10:29:51 +08:00
  • 1041f240c0 [NPU]grok2 model support (#17119) KnightLTC 2026-03-03 10:24:10 +08:00
  • 8b4c387aa2 [AMD] Update AITER Scout Workflow (#19735) YC Tseng 2026-03-03 05:14:31 +03:00
  • e6e02ec938 [diffusion]: Add model detectors and warning for quantized diffusion models (#18041) Ratish P 2026-03-03 07:16:25 +05:30
  • 145ae518ac [Diffusion] Revert 18619 (#19510) Xiaoyu Zhang 2026-03-03 08:15:15 +08:00
  • 6822941514 [FlashInfer] Bump FlashInfer version from 0.6.3 to 0.6.4 (#19005) Mohammad Miadh Angkad 2026-03-03 08:12:09 +08:00
  • 3f9fc8b848 [Qwen3.5] Fix missing quant_config in Qwen3VL (#19291) Mohammad Miadh Angkad 2026-03-03 06:07:51 +08:00
  • cc860a2198 [TestFix] change LoRA tests to use NVIDIA adapter instead of Nutanix (#19642) Glen Liu 2026-03-02 15:55:41 -05:00
  • 51ee17ce44 [diffusion] move skills dir (#19697) Xiaoyu Zhang 2026-03-03 02:51:29 +08:00
  • bdffb027a8 [CI] fix: handle missing repo in lora notebook (#19700) shuwenn 2026-03-03 02:27:32 +08:00
  • 05950853bc [Diffusion] [NPU] Add CI tests for FLUX (#19001) Makcum888e 2026-03-02 20:40:22 +03:00
  • c64274c746 Piecewise Cuda Graph set default (#16331) Yuwei An 2026-03-02 07:18:07 -08:00
  • 468e3dc56b [Qwen3.5] Set full attn_backend to trtllm_mha on SM100 by default when possible (#19030) hlu1 2026-03-02 07:14:53 -08:00
  • 2d183c4e6d [Feat] add PP Support for minimax-m2 series (#19577) 0xNullPath 2026-03-02 23:13:59 +08:00
  • 5833ea684d [diffusion] fix: make input/output file save paths configurable and disableable (#19580) Ruihang Li 2026-03-02 23:02:33 +08:00
  • 53de53fb53 [jit_kernel] Tiny unify jit_kernel tests style (#19694) Xiaoyu Zhang 2026-03-02 21:33:59 +08:00
  • 714c53d609 [NPU] support PD disaggregation on ascend when using PP (#14908) Hexq0210 2026-03-02 21:33:16 +08:00
  • eaf18ebe8d [sgl]add pin_mem to avoid cpu->gpu copy sync point (#19590) Bi Xue 2026-03-02 05:08:31 -08:00
  • b3718982a1 [Feature] add feature mla_ag_after_qlora for dsv3.2 (#19428) JiaruiChang5268 2026-03-02 20:00:31 +08:00
  • 3f36f27eae [Bugfix] Fix nixl and mori backend for missing decode tp size in PD module (#19690) Shangming Cai 2026-03-02 19:55:19 +08:00
  • 8df9b8dce9 [diffusion] fix: skip USP for cross-attention with replicated KV for wan (#19419) AichenF 2026-03-02 19:52:08 +08:00
  • da2a0240f7 Add GLM45 tool interruption support (#17714) Leoyzen 2026-03-02 19:34:12 +08:00
  • 7579ab3f33 Enhance error resilience in dump comparator (#19685) fzyzcjy 2026-03-02 19:08:35 +08:00
  • e5ef845cad Support multiple verbosity in dump comparator (#19684) fzyzcjy 2026-03-02 18:47:30 +08:00
  • 3dd4649b42 Beautify text output in dump comparator (#19683) fzyzcjy 2026-03-02 18:47:01 +08:00
  • 5bf3deb4bc Trace execution information in dump comparator (#19682) fzyzcjy 2026-03-02 18:46:27 +08:00
  • 3ebd85bf1c Enhance sglang engine dumping tests in dump comparator (#19681) fzyzcjy 2026-03-02 18:46:03 +08:00
  • abdc0ee71f Support directory detection in dump comparator (#19680) fzyzcjy 2026-03-02 18:45:35 +08:00
  • 6980416149 Support non orthogonal parallel axes and explicit replication annotation in dump comparator (#19679) fzyzcjy 2026-03-02 18:44:33 +08:00
  • a70dd11011 Support flattened dims in dump comparator (#19678) fzyzcjy 2026-03-02 18:43:01 +08:00
  • 15e83eea61 Enhance replication check, matching pattern, logging in dump comparator (#19677) fzyzcjy 2026-03-02 18:42:27 +08:00
  • ec44bc82ab Support presets and arbitrary skipping keys in dump comparator (#19676) fzyzcjy 2026-03-02 18:41:49 +08:00
  • 2e15c015c0 [diffusion] feat: Add --model-id for config resolution; deprecate model_detectors (#19607) Mick 2026-03-02 16:39:53 +08:00
  • f2c5503542 [AMD] AMD AITER Scout Workflow (#19467) YC Tseng 2026-03-02 11:10:44 +03:00
  • 98f47d8175 [AMD] Add Qwen3-Coder-Next accuracy and functionality test scripts for MI35x 8-GPU (#18608) jacky.cheng 2026-03-02 15:52:47 +08:00
  • 15af26d1e8 Add aiter attention support in prefill-attention-backend of gpt-oss (#18282) kk 2026-03-02 15:39:24 +08:00
  • f7da379b61 feat: TTL-based prefix pinning with refresh-on-hit for HiRadixCache (#18941) ishandhanani 2026-03-02 01:27:22 -06:00
  • e42fa009d4 [Diffusion] diffusion profile and opt skills (#19540) Xiaoyu Zhang 2026-03-02 15:06:29 +08:00
  • 07ef5f7be1 Remove sync points in mamba cache + prefill cudagraph plumbing for DP (#19639) Leon Gao 2026-03-01 23:03:42 -08:00
  • 4726073865 Fix mamba2 mixer ci test (#19658) Ke Bao 2026-03-02 14:55:15 +08:00
  • 922aad2faa Cleanup disagg decode prebuilt flow and add cross-stream sync in merge_batch (#19568) Baidu-AIAK 2026-03-02 13:52:27 +08:00
  • 0e53cee1f6 [CI] Disable test_lora_tp CUDA CI during H100 to H200 transition (#19654) sglang-bot 2026-03-01 21:43:32 -08:00
  • ec9775491f Add bisect ci claude code skill (#19649) Kangyan-Zhou 2026-03-01 20:53:40 -08:00
  • 57c5c343d7 [diffusion] model: support Hunyuan3D-2 (#18170) Prozac614 2026-03-02 12:28:05 +08:00
  • f6ee6dc8c3 [JIT-kernel] Add unit test for nsa indexer fused_store_k_cache (#19389) Yuan Luo 2026-03-02 12:18:11 +08:00
  • 0a6678bf3a [PD] Remove unused server args for disaggregation (#19618) Shangming Cai 2026-03-02 11:38:50 +08:00
  • e5edf222cd [WIP]enable mxfp8 on nvidia sm120 (#19112) Henry 2026-03-01 22:06:43 -05:00
  • e3e71f275a docs: refactor speculative decoding doc (#19186) shuwenn 2026-03-02 11:03:20 +08:00
  • 20282f5664 [fix typo] expert_indicies -> expert_indices (#19627) SoluMilken 2026-03-02 09:37:34 +08:00
  • f51ddba131 feat: add FA4 SM90 paged KV decode support & update attention docs (#18442) zwang86 2026-03-01 17:12:19 -08:00
  • 8a0b7575b0 [Test] add unit test for skipping already preempted request (#18912) Glen Liu 2026-03-01 18:44:15 -05:00