Commit Graph

6355 Commits

Author SHA1 Message Date
Simo Lin
5259becd3c [bug] fix router installation to include additional dependency (#12348) 2025-10-29 12:45:18 -07:00
AichenF
ed1044ac1b support cutlass fp4 kernel in sm120 (#11737) 2025-10-29 12:25:16 -07:00
Chang Su
d717e73ea8 [router] refactor mcp to use LRU and fix pooling bug (#12346) 2025-10-29 12:18:54 -07:00
weiliang
a18161875c Fix Flashinfer Backend for SM120 Usage (#12325) 2025-10-29 11:51:55 -07:00
Minglei Zhu
e39628fd07 [2/2] Deepseek deterministic: support deepseek v3 deterministic inference on 8 x H200 (#12095) 2025-10-29 11:49:04 -07:00
b8zhong
bacb3825fe fix: llama 4 + trtllm gen + fp8 kv cache incompatibility (#12347) 2025-10-29 11:31:02 -07:00
Simo Lin
b53d9e11c6 [bug] fix router pypi license file (#12345) 2025-10-29 10:05:00 -07:00
zejunchen-zejun
8a6838212a [Fix] fix type issue of env flag value MODELOPT_MAX_TOKENS_PER_EXPERT (#11709)
Signed-off-by: zejunchen-zejun <zejun.chen@amd.com>
2025-10-29 09:44:05 -07:00
Xiaoyu Zhang
52694b60da Triton fused_moe_kernel support ep moe tuning (#12343) 2025-10-29 23:16:09 +08:00
Chang Su
400bddf24c [router] fix router release workflow and add build test in PR (#12315) 2025-10-29 08:11:45 -07:00
Feng Su
1e90fe2ed1 [Bug fix] trace: fix import error in mini_lb if sgl-router image does not install sglang (#12338) 2025-10-29 07:45:08 -07:00
Yuhong Guo
caa5d2967c feat: return partial generation results when aborting requests in waiting queue (#11673) 2025-10-29 22:03:00 +08:00
Rain H
750940ae36 Eagle3 DP attention for Qwen3 MoE (#12002) 2025-10-29 20:25:17 +08:00
Liangsheng Yin
42f8ea4030 [Test] Fix session control test (#12336) 2025-10-29 18:28:04 +08:00
Liangsheng Yin
14cbe42fd3 Refactor abortion in event loop (#12312) 2025-10-29 18:25:20 +08:00
Baizhou Zhang
685c06451f [ci] Try fixing broken CIs (#12317) 2025-10-29 01:13:51 -07:00
Liana Koleva
1357397a34 feat: preview filename from tuning_fused_moe_triton.py (#12276)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
2025-10-29 16:12:25 +08:00
hlu1
42e1a72efb [Deepseek V3.2] Enable flashmla_auto with MTP (#12294)
Signed-off-by: Hao Lu <14827759+hlu1@users.noreply.github.com>
2025-10-28 23:51:20 -07:00
b8zhong
83a7c89c3f followup fix for llama 4 trtllm flashinfer backend (#12314) 2025-10-28 22:17:08 -07:00
Yuzhen Zhou
0380ca82ef Add Batch‑Invariant RMSNorm (#12144) 2025-10-28 21:05:57 -07:00
Yingchun Lai
ec92b0cefe EPLB: prefer to use physical experts in the same gpu or node (#10874) 2025-10-28 21:01:11 -07:00
Liana Koleva
e03b6beeb1 doc: improve modelopt error description (#12269) 2025-10-28 20:57:55 -07:00
Yingchun Lai
5e36a0b455 [metrics][EPLB]: Support selected count of physical experts on each GPU (#9825) 2025-10-28 20:56:19 -07:00
Gao016
0297773a2f a tiny fix for support deepseek bf16 weights (#12313)
Co-authored-by: gaochang <gaochang@U-19PX2WQ1-0350.local>
2025-10-28 20:46:44 -07:00
Baizhou Zhang
587deb15a7 [hotfix] Fix pytest not found in CI (#12311) 2025-10-29 11:07:36 +08:00
Cheng Wan
83087247d1 [hotfix] missing w13_weight_fp8 and w2_weight_fp8 in UE8M0 requantization (#12259) 2025-10-28 19:10:38 -07:00
Xiaoyu Zhang
334543ff3b Add continuous_usage_stats support for streaming responses (#12241) 2025-10-29 10:01:23 +08:00
b8zhong
c143f416ce fix: Llama 4 BF16 load on Blackwell (#12308) 2025-10-28 18:59:01 -07:00
Chang Su
b48354c537 [router][grpc] Fix inconsistent behavior of conversation_id not found (#12299) 2025-10-28 16:24:43 -07:00
fzyzcjy
29195aaa6e Super tiny fix expert distribution dump error (#12271) 2025-10-28 15:20:55 -07:00
hlu1
0ee831dee0 Update deepseek_v32.md (#12296) 2025-10-28 14:52:38 -07:00
bmac3
8d6ab1cb88 fix seqlen bug for trtllm_mla's draft_extend (#12295) 2025-10-28 14:47:47 -07:00
Simo Lin
84a9d0eab2 [router] support arm, windows, mac, linux, reduce wheel size and number (#12285) 2025-10-28 12:18:51 -07:00
Keyang Ru
737b58d6dc [rust][ci] Add end-to-end tests for Oracle history backend (#12233) 2025-10-28 10:57:32 -07:00
b8zhong
77225d602a Use Flashinfer TRT-LLM as Llama 4 compatible MoE backend (#11928) 2025-10-28 10:39:43 -07:00
ybyang
9c6e25d2a6 doc for logit_bias (#12188) 2025-10-28 10:32:12 -07:00
fzyzcjy
2a3763c335 Tiny fix sgl-kernel related CI installing the wrong binary (#12283) 2025-10-28 10:29:06 -07:00
Trevor Morris
fdd00295b5 Fix 'BypassedTopKOutput' object has no attribute 'topk_weights' for DeepEP (#12231) 2025-10-28 09:28:25 -07:00
Simo Lin
25e73640f4 [router] upgrade grpc dependency and py 3.13 3.14 support (#12284) 2025-10-28 08:51:32 -07:00
sogalin
0da9845ef1 Modify rocm.Dockerfile (#12274) 2025-10-28 08:36:32 -07:00
Keyang Ru
9288544180 [router] Fix type unmatch during validation (#12257) 2025-10-28 06:04:54 -07:00
Yineng Zhang
64cf868eba chore: cleanup quant deps (#12268) 2025-10-28 02:03:57 -07:00
Yineng Zhang
ea39952797 Revert "[Feature] PD-Multiplexing Context and Scheduler." (#12267) 2025-10-28 02:00:37 -07:00
Shangming Cai
41a113356a Fix potential eos bug on decode instance when PD is enabled (#12206)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2025-10-28 01:29:02 -07:00
Xuchun Shang
a1f2dc90e4 [Bug fix] [PP] fix wrong dtype for quantified model (#12247)
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
2025-10-28 01:27:24 -07:00
Feng Su
ea96106000 [Feature] Sglang Tracing: Fine-Grained Tracking for Request Latency - Part 2 (#10804)
Signed-off-by: Feng Su <sufeng@linux.alibaba.com>
2025-10-28 01:25:46 -07:00
Cheng Wan
b1e13e7cea [hotfix] Incorrect CombineOverlapArgs in SBO (#12230) 2025-10-28 01:23:06 -07:00
Chenxi Li
cc7b04a29c Feature/Add GET endpoint to query loaded LoRA adapters (#12229) 2025-10-28 01:22:00 -07:00
Simo Lin
d85d6dba3b [router] configure workflow retries and timeout based on routerConfig (#12252) 2025-10-28 00:41:20 -07:00
Simo Lin
c5642a7a7a [router] use mcp struct from sdk and clean up code across codebase (#12249) 2025-10-28 00:33:10 -07:00