Commit Graph

  • d6fee73d1f Support nvidia/NVIDIA-Nemotron-Nano-9B-v2-FP8/NVFP4 (#11866) Netanel Haber 2025-10-23 12:29:02 +03:00
  • 36a4cad7b0 Support overlap-spec-v2 with trtllm_mla attention backend (#11821) Qiaolin Yu 2025-10-23 01:55:35 -07:00
  • 65d376b491 aiter update to v0.1.6.post1 (#12004) HAI 2025-10-22 23:53:05 -07:00
  • c23eda8589 Fix incorrect KV indices creation when page_size=32 in TRTLLM MLA backend (#11985) yinghui 2025-10-22 22:44:45 -07:00
  • 138ff23187 Allow to disable batch decoding. (#11944) Jue WANG 2025-10-22 21:57:12 -07:00
  • 13fb8b5489 [CPU] Optimize FP16 decode_attention_cpu (#10652) blzheng 2025-10-23 12:39:51 +08:00
  • 81fd2b0ee0 fix(deepep): resolve benchmark failure on 4×IB-card setup by aligning tuning config with DeepEP commit bdd119f8 (#11965) Zhengyi Lai 2025-10-23 12:20:54 +08:00
  • 007b849b0e [CPU] misc updates (#11906) Zaili Wang 2025-10-23 12:10:05 +08:00
  • 8612811d85 Bump grace blackwell DeepEP version (#11990) fzyzcjy 2025-10-23 12:08:12 +08:00
  • e7aa4664b3 [NVIDIA] Build CUDA 13 (#11299) Johnny 2025-10-22 20:03:12 -07:00
  • 4d4feccbb2 [ROCm] Remove vLLM rope dependency & use AITER impl (#11322) b8zhong 2025-10-22 19:17:34 -07:00
  • 99c92ff24b [AMD] Support a new flag to disable quant on parallelLinear layer if required (#11811) jacky.cheng 2025-10-23 10:16:15 +08:00
  • 6ade6a02d4 [grpc] Support gRPC standard health check (#11955) Chang Su 2025-10-22 16:59:09 -07:00
  • 983ef22cf3 [Doc] Update deterministic inference flag in server_arguments.md (#11978) Baizhou Zhang 2025-10-22 16:12:15 -05:00
  • 164302c7df Implement BGE-M3 Sparse Embeddings in SGLang (#10869) Christian Bahls 2025-10-22 22:46:16 +02:00
  • 5dccf69713 [router] create worker removal step and clean up worker manager (#11921) Simo Lin 2025-10-22 13:26:06 -07:00
  • eec9e471ca [NVIDIA] Update to leverage flashinfer trtllm FP4 MOE throughput kernel (#11563) jiahanc 2025-10-22 13:11:16 -07:00
  • 6d535b719f Revert "Recapture cuda graph after model weight update to resolve IMA error " (#11980) Lianmin Zheng 2025-10-22 11:50:26 -07:00
  • fdcb1d13c5 [BUG] AttributeError: 'DeepEPMoE' object has no attribute 'use_w4a… (#11977) yuho 2025-10-23 02:29:55 +08:00
  • d7e834d6ba [6/n]decouple quantization implementation from vLLM dependency (#10750) Hongbo Xu 2025-10-23 02:07:55 +08:00
  • 200a3c0bb1 [Documentation] add doc for deterministic inference (#11956) Minglei Zhu 2025-10-22 10:36:15 -07:00
  • 77258ce039 [router] Support multiple worker URLs for OpenAI router (#11723) Keyang Ru 2025-10-22 09:27:58 -07:00
  • 1d097aac87 [Fix] Remove unused import from triton_kernels_moe.py (#11967) Fan Yin 2025-10-22 21:02:57 +08:00
  • 7fceeef599 Fix flaky hicache test with mooncake backend (#11953) Shangming Cai 2025-10-22 21:00:47 +08:00
  • 88568c01eb [model] Support POINTSV15Chat (#9651) 996_icu 2025-10-22 16:58:17 +08:00
  • 904655c5fd [2/N] Added the core structure of elastic EP and the eplb algorithm with faulty rank (#10606) Hank Han 2025-10-22 16:13:31 +08:00
  • e028af6998 Fix mooncake dispatcher (#11908) Xun Sun 2025-10-22 16:11:49 +08:00
  • 80b2b3207a Enable native ModelOpt quantization support (3/3) (#10154) Zhiyu 2025-10-21 21:44:29 -07:00
  • 4b65ed42cc [NVIDIA] upstream FA4 and fix cccl path (#11929) Johnny 2025-10-21 21:18:25 -07:00
  • 23afdfd1c2 [sgl-kernel] support flashmla libtorch (#11717) Fan Yin 2025-10-22 12:17:50 +08:00
  • 9d61205dac [lint] improve ruff check (#11922) Liangsheng Yin 2025-10-22 11:32:50 +08:00
  • 590bc4b7a7 [router][grpc] Fix background tasks stored with wrong id (#11945) Chang Su 2025-10-21 18:38:51 -07:00
  • 63cfe1b032 [router] Add gRPC E2E test suite (#11790) Keyang Ru 2025-10-21 17:51:21 -07:00
  • 70f6309cd4 [router][grpc] Support v1/responses API (#11926) Chang Su 2025-10-21 17:41:48 -07:00
  • 704160017d fix: resolve flashinfer 0.4.1 import (#11940) Yineng Zhang 2025-10-21 17:19:57 -07:00
  • 87a92e459a Fix openai input_text type compatibility (#11935) Keyang Ru 2025-10-21 16:10:35 -07:00
  • c461e7714d [Auto Sync] Update forward_batch_info.py (20251021) (#11934) Yineng Zhang 2025-10-21 15:52:15 -07:00
  • fde2decf8b [BugFix][Qwen3-VL]: add metadata for video in qwen3-vl (#11377) Zheng Wengang 2025-10-22 06:36:01 +08:00
  • 9792b9d7e3 chore: upgrade flashinfer 0.4.1 (#11933) Yineng Zhang 2025-10-21 14:46:31 -07:00
  • ef4a8097b8 Rename flashmla kernel options of nsa backend for better readability (#11876) Baizhou Zhang 2025-10-21 15:14:16 -05:00
  • ebff4ee648 Update sgl-kernel and remove fast hadamard depedency (#11844) Baizhou Zhang 2025-10-21 15:13:54 -05:00
  • 2b1da821b5 [NVIDIA] Add new SMs support for Spark & Thor (#11287) Serge Panev 2025-10-21 11:02:24 -07:00
  • 97710ccd1a Fix flush cache API for spec v2 (#11918) Liangsheng Yin 2025-10-21 23:01:16 +08:00
  • f3cd5d2510 [CI] Fix b200 flashinfer installation (#11915) Shangming Cai 2025-10-21 22:28:50 +08:00
  • c61b0b294c [quantization][MoE] fix the check for tp_size / moe_ep_size / moe_intermediate_size / weight_block_size_n (#11702) Kai-Hsun Chen 2025-10-21 06:25:28 -07:00
  • e8640ee9be [smol] [perf] Inverse perm improvement (#11482) Vincent Zhong 2025-10-21 07:18:10 -04:00
  • d0a64c7e2c vlm: enforce pybase64 for image and str encode/decode (#10700) b8zhong 2025-10-21 04:05:32 -07:00
  • 05d3667ab9 [CI] disable glm4.1v and fix the flashinfer installation (#11902) Shangming Cai 2025-10-21 18:38:35 +08:00
  • 260fe755b6 Simplify multi-tokenizer (#11295) Zhengke Zhou 2025-10-21 16:33:29 +08:00
  • dbb16bedd5 Support Thinking Budget (via custom_logit_processor for OpenAI API) [Fix #6572] (#11416) ybyang 2025-10-21 16:27:56 +08:00
  • c1e1600373 [fix] fix ci uv install dependency (#11895) Hank Han 2025-10-21 16:23:34 +08:00
  • 852c0578fd [FEATURE] Add OpenAI-Compatible LoRA Adapter Selection (#11570) Neelabh Sinha 2025-10-21 00:44:33 -07:00
  • 7e6191c098 init support for KTransformers Heterogeneous Computing (#11487) Atream 2025-10-21 15:17:02 +08:00
  • 6f9b66bdda [AMD] Update wave-lang to 3.8.0 (#11878) Gaurav Verma 2025-10-20 23:11:09 -07:00
  • 8a801ee38d [router] release router 0.2.1 (#11885) Simo Lin 2025-10-20 21:08:45 -07:00
  • d9a20fd28a Use trtllm_mla decode kernel for draft extend in speculative decoding (#11664) Qiaolin Yu 2025-10-20 20:42:09 -07:00
  • b113c72e7a Init attention backend for Intel XPU (#10656) Meng, Hengyu 2025-10-21 11:41:28 +08:00
  • fb6cc7b000 Fix RotaryEmbedding for fp32 input (#11843) zhangdonghao-zdh 2025-10-21 10:56:48 +08:00
  • 8374a96e49 piecewise cuda graph support qwen3-moe (#11845) Xiaoyu Zhang 2025-10-21 10:55:49 +08:00
  • 74de76c685 Revise MRotaryEmbedding's forward (#11859) Yuan Luo 2025-10-21 10:38:29 +08:00
  • 9c0b1eb5ad [router][grpc] Fix wram-up random token ids for small models (#11887) Chang Su 2025-10-20 19:22:17 -07:00
  • 01f14a7ad2 [code move] move pp into a separate mixin (#11838) Lianmin Zheng 2025-10-20 18:46:56 -07:00
  • 1111030395 [router] clean up workflow logs to debug for implementation details logs (#11886) Simo Lin 2025-10-20 18:24:55 -07:00
  • 28ddfb37d7 fix(sql-router): fix conflict port in test (#11826) Tien Nguyen 2025-10-21 08:06:34 +07:00
  • e69094df64 [router][grpc] Remove continue_final_message in ChatTemplateParams and add minijinja-contrib (#11882) Chang Su 2025-10-20 18:03:09 -07:00
  • 43ad05907c [Auto Sync] Update scheduler.py, server_args.py (20251020) (#11875) Lianmin Zheng 2025-10-20 17:41:19 -07:00
  • b4948512b8 [router] remove encoding header for oai router (#11881) Simo Lin 2025-10-20 17:39:00 -07:00
  • ddcba74b4d [router] Worker Management Workflow Engine (#11868) Simo Lin 2025-10-20 17:00:22 -07:00
  • 0917c5da8c Support mixing cutedsl and deepgemm backend (#11807) fzyzcjy 2025-10-21 07:38:35 +08:00
  • 184a4df697 Replace function call with set literal (#11867) penguin_wwy 2025-10-21 01:39:16 +08:00
  • f7b1d8c5ab Fix acc len and gen throughput metrics when enabling overlap-spec (#11823) Qiaolin Yu 2025-10-20 10:34:38 -07:00
  • bfc3b3f786 [9/N] MoE Refactor: cleanup dispatcher interfaces (#11847) Cheng Wan 2025-10-20 10:11:46 -07:00
  • da5bde4d16 Tiny fix main lint (#11862) Liangsheng Yin 2025-10-20 19:57:24 +08:00
  • 276e7b3e4e [Feature] New structural tag support (#10691) DarkSharpness 2025-10-20 18:25:58 +08:00
  • 296f689242 fix(server_args): handle tokenizer init conflicts (#11776) ishandhanani 2025-10-20 00:27:19 -07:00
  • 9edb7b5123 [AMD CI] Populate image cache in nightly docker release. (#11822) Sai Enduri 2025-10-20 00:04:04 -07:00
  • e53bf44243 Update amd gpu install docs. (#11849) Sai Enduri 2025-10-20 00:03:26 -07:00
  • d383e6616e [Model] Add Olmo 3 model support (#11396) Shane A 2025-10-19 23:59:16 -07:00
  • 984fbeb16b Revert "[CI Monitor] Ci monitor only deal with main branch in default" (#11846) Xiaoyu Zhang 2025-10-20 13:06:40 +08:00
  • a2ba0bc3df Tiny clean up for PD module and doc (#11747) Shangming Cai 2025-10-20 11:52:42 +08:00
  • 6d2d0ce285 [PD] Improve eagle acceptance rate by transferring draft model hidden states (#10801) Ziming Huang 2025-10-20 11:52:18 +08:00
  • 271d3d0d50 Support mrope triton kernel and add unit test (#11722) Yuan Luo 2025-10-20 11:51:07 +08:00
  • c4e81e64fb [Feature] Use current greenctx stream to communicate in PD-Multiplexing. (#11594) ykcombat 2025-10-20 10:58:20 +08:00
  • c726d44cc7 Recapture cuda graph after model weight update to resolve IMA error (#11780) harrisonlimh 2025-10-19 19:50:03 -07:00
  • 283c8ba031 chore: bump sgl-kernel version to 0.3.16.post3 (#11733) sglang-bot 2025-10-19 19:44:15 -07:00
  • cae3956585 check master server for mooncake store (#10510) huangtingwei 2025-10-20 09:37:09 +08:00
  • 27a223aba4 Improve Kernel Build Time (#11508) Kangyan-Zhou 2025-10-19 18:11:48 -07:00
  • 53529f46cc Fix version bump script to handle TOML files with outdated versions (#11787) Kangyan-Zhou 2025-10-19 18:10:26 -07:00
  • 24ed3f32c0 fix(ci): Fix CI Monitor limit parameter and add CI Analysis to summary (#11832) Xiaoyu Zhang 2025-10-20 09:08:34 +08:00
  • 44f0ece9fc [Doc] Update documents for FA4 (#11778) Baizhou Zhang 2025-10-19 19:40:38 -05:00
  • be0058bc05 [BugFix] replace the input_to_float8 used in dsv2 (#11612) Liu-congo 2025-10-20 08:34:13 +08:00
  • 9e3be1fa2a Tiny bump DeepEP version in ARM blackwell (#11810) fzyzcjy 2025-10-20 08:15:14 +08:00
  • a8ba32798e Fix triton_kernels import error on some hardwares (#11831) fzyzcjy 2025-10-20 08:14:47 +08:00
  • 3b80232d06 [DeepseekV32] Add fast_topk_transform_ragged_fused kernel (#11815) hlu1 2025-10-19 17:13:39 -07:00
  • 252dc4e112 [NVIDIA] FA3/FA4 Fix (#11606) Johnny 2025-10-20 02:10:10 +02:00
  • cbb5fc2edc [CI] Add CI test for DeepSeek V3.2 MTP (#11835) Baizhou Zhang 2025-10-19 19:00:25 -05:00
  • 53fb229f53 [logprobs] Enable local deterministic logrprobs testing with strict threshold (#10994) Night 2025-10-19 13:30:39 -07:00
  • 4fff1ec1d9 Deterministic Mode: Add 1-stage triton kernel for prefill (#11147) Stefan He 2025-10-19 10:47:36 -07:00
  • 7a020e0f3b [Test] Add basic matched stop for beta eagle (#11833) Liangsheng Yin 2025-10-20 01:17:00 +08:00
  • 48738af7f9 [CI] always print back trace in retry() (#11834) Liangsheng Yin 2025-10-20 01:12:49 +08:00