Commit Graph

  • db15148c8a [router] update router docker to use maturin and build from local (#12350) Chang Su 2025-10-29 16:00:57 -04:00
  • 5259becd3c [bug] fix router installation to include additional dependency (#12348) Simo Lin 2025-10-29 12:45:18 -07:00
  • ed1044ac1b support cutlass fp4 kernel in sm120 (#11737) AichenF 2025-10-30 03:25:16 +08:00
  • d717e73ea8 [router] refactor mcp to use LRU and fix pooling bug (#12346) Chang Su 2025-10-29 15:18:54 -04:00
  • a18161875c Fix Flashinfer Backend for SM120 Usage (#12325) weiliang 2025-10-30 02:51:55 +08:00
  • e39628fd07 [2/2] Deepseek deterministic: support deepseek v3 deterministic inference on 8 x H200 (#12095) Minglei Zhu 2025-10-29 11:49:04 -07:00
  • bacb3825fe fix: llama 4 + trtllm gen + fp8 kv cache incompatibility (#12347) b8zhong 2025-10-29 11:31:02 -07:00
  • b53d9e11c6 [bug] fix router pypi license file (#12345) Simo Lin 2025-10-29 10:05:00 -07:00
  • 8a6838212a [Fix] fix type issue of env flag value MODELOPT_MAX_TOKENS_PER_EXPERT (#11709) zejunchen-zejun 2025-10-30 00:44:05 +08:00
  • 52694b60da Triton fused_moe_kernel support ep moe tuning (#12343) Xiaoyu Zhang 2025-10-29 23:16:09 +08:00
  • 400bddf24c [router] fix router release workflow and add build test in PR (#12315) Chang Su 2025-10-29 11:11:45 -04:00
  • 1e90fe2ed1 [Bug fix] trace: fix import error in mini_lb if sgl-router image does not install sglang (#12338) Feng Su 2025-10-29 22:45:08 +08:00
  • caa5d2967c feat: return partial generation results when aborting requests in waiting queue (#11673) Yuhong Guo 2025-10-29 22:03:00 +08:00
  • 750940ae36 Eagle3 DP attention for Qwen3 MoE (#12002) Rain H 2025-10-29 20:25:17 +08:00
  • 42f8ea4030 [Test] Fix session control test (#12336) Liangsheng Yin 2025-10-29 18:28:04 +08:00
  • 14cbe42fd3 Refactor abortion in event loop (#12312) Liangsheng Yin 2025-10-29 18:25:20 +08:00
  • 685c06451f [ci] Try fixing broken CIs (#12317) Baizhou Zhang 2025-10-29 01:13:51 -07:00
  • 1357397a34 feat: preview filename from tuning_fused_moe_triton.py (#12276) Liana Koleva 2025-10-29 01:12:25 -07:00
  • 42e1a72efb [Deepseek V3.2] Enable flashmla_auto with MTP (#12294) hlu1 2025-10-28 23:51:20 -07:00
  • 83a7c89c3f followup fix for llama 4 trtllm flashinfer backend (#12314) b8zhong 2025-10-28 22:17:08 -07:00
  • 0380ca82ef Add Batch‑Invariant RMSNorm (#12144) Yuzhen Zhou 2025-10-29 00:05:57 -04:00
  • ec92b0cefe EPLB: prefer to use physical experts in the same gpu or node (#10874) Yingchun Lai 2025-10-29 12:01:11 +08:00
  • e03b6beeb1 doc: improve modelopt error description (#12269) Liana Koleva 2025-10-28 20:57:55 -07:00
  • 5e36a0b455 [metrics][EPLB]: Support selected count of physical experts on each GPU (#9825) Yingchun Lai 2025-10-29 11:56:19 +08:00
  • 0297773a2f a tiny fix for support deepseek bf16 weights (#12313) Gao016 2025-10-29 11:46:44 +08:00
  • 587deb15a7 [hotfix] Fix pytest not found in CI (#12311) Baizhou Zhang 2025-10-28 20:07:36 -07:00
  • 83087247d1 [hotfix] missing w13_weight_fp8 and w2_weight_fp8 in UE8M0 requantization (#12259) Cheng Wan 2025-10-28 19:10:38 -07:00
  • 334543ff3b Add continuous_usage_stats support for streaming responses (#12241) Xiaoyu Zhang 2025-10-29 10:01:23 +08:00
  • c143f416ce fix: Llama 4 BF16 load on Blackwell (#12308) b8zhong 2025-10-28 18:59:01 -07:00
  • b48354c537 [router][grpc] Fix inconsistent behavior of conversation_id not found (#12299) Chang Su 2025-10-28 19:24:43 -04:00
  • 29195aaa6e Super tiny fix expert distribution dump error (#12271) fzyzcjy 2025-10-29 06:20:55 +08:00
  • 0ee831dee0 Update deepseek_v32.md (#12296) hlu1 2025-10-28 14:52:38 -07:00
  • 8d6ab1cb88 fix seqlen bug for trtllm_mla's draft_extend (#12295) bmac3 2025-10-28 14:47:47 -07:00
  • 84a9d0eab2 [router] support arm, windows, mac, linux, reduce wheel size and number (#12285) Simo Lin 2025-10-28 12:18:51 -07:00
  • 737b58d6dc [rust][ci] Add end-to-end tests for Oracle history backend (#12233) Keyang Ru 2025-10-28 10:57:32 -07:00
  • 77225d602a Use Flashinfer TRT-LLM as Llama 4 compatible MoE backend (#11928) b8zhong 2025-10-28 10:39:43 -07:00
  • 9c6e25d2a6 doc for logit_bias (#12188) ybyang 2025-10-29 01:32:12 +08:00
  • 2a3763c335 Tiny fix sgl-kernel related CI installing the wrong binary (#12283) fzyzcjy 2025-10-29 01:29:06 +08:00
  • fdd00295b5 Fix 'BypassedTopKOutput' object has no attribute 'topk_weights' for DeepEP (#12231) Trevor Morris 2025-10-28 09:28:25 -07:00
  • 25e73640f4 [router] upgrade grpc dependency and py 3.13 3.14 support (#12284) Simo Lin 2025-10-28 08:51:32 -07:00
  • 0da9845ef1 Modify rocm.Dockerfile (#12274) sogalin 2025-10-28 23:36:32 +08:00
  • 9288544180 [router] Fix type unmatch during validation (#12257) Keyang Ru 2025-10-28 06:04:54 -07:00
  • 64cf868eba chore: cleanup quant deps (#12268) Yineng Zhang 2025-10-28 02:03:57 -07:00
  • ea39952797 Revert "[Feature] PD-Multiplexing Context and Scheduler." (#12267) Yineng Zhang 2025-10-28 02:00:37 -07:00
  • 41a113356a Fix potential eos bug on decode instance when PD is enabled (#12206) Shangming Cai 2025-10-28 16:29:02 +08:00
  • a1f2dc90e4 [Bug fix] [PP] fix wrong dtype for quantified model (#12247) Xuchun Shang 2025-10-28 16:27:24 +08:00
  • ea96106000 [Feature] Sglang Tracing: Fine-Grained Tracking for Request Latency - Part 2 (#10804) Feng Su 2025-10-28 16:25:46 +08:00
  • b1e13e7cea [hotfix] Incorrect CombineOverlapArgs in SBO (#12230) Cheng Wan 2025-10-28 01:23:06 -07:00
  • cc7b04a29c Feature/Add GET endpoint to query loaded LoRA adapters (#12229) Chenxi Li 2025-10-28 01:22:00 -07:00
  • d85d6dba3b [router] configure workflow retries and timeout based on routerConfig (#12252) Simo Lin 2025-10-28 00:41:20 -07:00
  • c5642a7a7a [router] use mcp struct from sdk and clean up code across codebase (#12249) Simo Lin 2025-10-28 00:33:10 -07:00
  • 691c8534cf Support releasing CUDA graph memory when paused (#7873) fzyzcjy 2025-10-28 14:40:50 +08:00
  • d2b8c4123e Opt fused triton moe: add tma for down proj kernel (#10567) Yongfei Xu 2025-10-28 14:26:17 +08:00
  • bf8f7a944f Add per-request retraction count (#11177) Scott Lee 2025-10-27 23:22:34 -07:00
  • 81a632ace6 [DeepseekV32] Enable flashmla_prefill kernel with fp8 kvcache (#11655) hlu1 2025-10-27 23:11:48 -07:00
  • 83b2240074 [router] remove code duplication (#12245) Simo Lin 2025-10-27 23:01:44 -07:00
  • 285a8e6986 docker: add CUDA13 support in dockerfile and update GDRCopy/NVSHMEM for blackwell support (#11517) ishandhanani 2025-10-27 22:00:54 -07:00
  • 813bd6f85c [2/2] Use moe_sum_reduce cuda kernel (#10654) Yuan Luo 2025-10-28 12:01:57 +08:00
  • 729f612dc6 Update openai package version to 2.6.1 (#12222) Xinyuan Tong 2025-10-27 20:23:40 -07:00
  • 899453ac50 Use explicit uint64 dtype for Tensor data_ptr() to avoid overflow (#11994) jianan-gu 2025-10-28 10:05:57 +08:00
  • ce832d7034 Add env var to control custom Triton kernel cache and set CSGMV as default backend. (#12176) Lifu Huang 2025-10-27 17:49:32 -07:00
  • 88596739a4 Support running FP4 Deepseek on SM120. (#11708) weiliang 2025-10-28 08:37:49 +08:00
  • a6ea3add76 [Auto Sync] Update scheduler.py, spec_info.py, run_suite.py... (20251027) (#12235) Yineng Zhang 2025-10-27 17:21:08 -07:00
  • 326c84c493 Compiling rope while preserving true on policy (#12161) fzyzcjy 2025-10-28 08:02:17 +08:00
  • 8da608cce0 fix: AttributeError: 'NixlKVManager' object has no attribute 'prefill_tp_size_table' (#12234) gongwei-130 2025-10-27 15:57:33 -07:00
  • 9fc3e8aac7 Add support for Matryoshka embeddings (#126) (#11142) satyamk7054 2025-10-27 11:49:36 -07:00
  • c11b34d599 rope xpu: fix missing argument 'fused_set_kv_buffer_arg' and replace native with sgl_kernel_xpu impl (#12006) Chunyuan WU 2025-10-28 01:02:18 +08:00
  • 05ad28f25e [Feature] PD-Multiplexing Context and Scheduler. (#11592) ykcombat 2025-10-28 00:54:43 +08:00
  • 0cae873fcd check_offload_progress more frequently (#11656) pansicheng 2025-10-28 00:37:38 +08:00
  • a8b91f6b2d improve mimax-m2 rmsnorm precision (#12186) Haichao Zhu 2025-10-28 00:01:42 +08:00
  • 959d1ab84b fix(metrics): double times add_latency for DECODE_BOOTSTRAP (#12209) Jimmy 2025-10-27 23:59:48 +08:00
  • 6c1c193308 [Detokenizer Manager] Cleanup state when reqs are finished (#12205) Muqi Li 2025-10-27 23:59:09 +08:00
  • 3029d30189 Fix crash after flush cache (#12107) cctry 2025-10-27 08:52:27 -07:00
  • f389f01714 Optimize triton_mrope with torch compile (#12112) Yuan Luo 2025-10-27 23:49:22 +08:00
  • caa4819bfc Add support for AutoRound quantized models (#10153) Weiwei 2025-10-27 18:17:29 +08:00
  • a88b006ecf GLM-4-0414 and GLM-4.1V Code Refactor (#12117) Yuxuan Zhang 2025-10-27 16:57:07 +08:00
  • ce112c07fe [sgl-kernel][4/N]Support Expert Specialization Grouped GEMM (#12080) Qi Yuhang 2025-10-27 16:20:01 +08:00
  • f7dc2f334b [sgl-kernel] feat: Support sm120 cutlass fp8 gemm kernel (#9403) Joonchen Liau 2025-10-27 14:45:45 +08:00
  • cd784faf20 docs: update contact (#12192) Yineng Zhang 2025-10-26 23:23:39 -07:00
  • 75c09e1ffe [Fix] Fix cu130 sgl-kernel wheel renaming (#12173) Baizhou Zhang 2025-10-26 22:44:05 -07:00
  • 09af0a7b5a [sgl-route] Optimize the use of constant slices and retain to simplif… (#12159) rongfu.leng 2025-10-27 12:14:33 +08:00
  • c8d385ce68 [doc] add example of using w4fp8 for Deepseek (#12057) Kevin_Xiong 2025-10-27 11:10:38 +08:00
  • 55d75e11bd chore: bump SGLang version to 0.5.4.post1 (#12169) sglang-bot 2025-10-27 09:35:20 +08:00
  • 3f4cc0aff0 Remove description for --enable-beta-spec argument (#12177) Xinyuan Tong 2025-10-26 16:49:39 -07:00
  • cadfae666d fix broken deepep/flashmla install in container by adding --no-build-isolation (#12170) ishandhanani 2025-10-26 16:39:51 -07:00
  • da1766e444 Remove deprecated --enable-beta-spec argument and fix b200 test (#12167) Kangyan-Zhou 2025-10-26 15:54:40 -07:00
  • a124b517f6 [router] Remove SharedXxxStorage type aliases to make Arc explicit (#12171) Chang Su 2025-10-26 14:48:07 -07:00
  • d05a968ba8 [router][grpc] Add ResponsesContext and fix error propagation in responses api (#12164) Chang Su 2025-10-26 14:07:19 -07:00
  • 94aad0de99 [misc][grpc] Remove duplicate log (#12168) Chang Su 2025-10-26 14:06:59 -07:00
  • 7ebc28f5d6 [WIP] support MiniMax M2 model (#12129) 赵晨阳 2025-10-26 13:58:54 -07:00
  • b89111d69b add gitignore for claude code and serena mcp (#12166) Simo Lin 2025-10-26 13:30:17 -07:00
  • 0b3b3e9a69 transfer mrope_position_delta to device when first running (#11047) ash-sigh 2025-10-27 04:06:09 +08:00
  • a1d5bc4cce Avoid using flashinfer_allreduce_fusion when dp attention is enabled. (#11632) Elfie Guo 2025-10-26 12:31:14 -07:00
  • a8023891f6 model: support NVILA and NVILA Lite (#10399) Zijian Zhang 2025-10-27 00:58:09 +08:00
  • 0103f374ba Support DeepGEMM for deterministic inference (#12142) fzyzcjy 2025-10-26 22:36:17 +08:00
  • 96a5a949f6 [Fix] fix allreduce bug in Piecewise Graph (#12106) zyksir 2025-10-26 21:15:48 +08:00
  • ea385ae85a Fix ITL metrics when using openai endpoint with spec (#12156) Liangsheng Yin 2025-10-26 18:06:25 +08:00
  • 9e949e5877 [router] centralize mcp tool args handling (#12155) Simo Lin 2025-10-26 01:47:12 -07:00
  • 6dbb569bc3 [router][grpc] Fix tool call id in parse_json_schema_response (#12152) Chang Su 2025-10-26 01:40:53 -07:00
  • 5994e6c373 Do not use MagicMock to mock server_args in tests (#12154) Liangsheng Yin 2025-10-26 16:00:46 +08:00