Commit Graph

  • e5e65e3d2a [router][grpc] Support vllm backend for grpc router (#13120) Chang Su 2025-11-12 02:29:20 -08:00
  • ffeb28ba6f fix: duplicate resize images logic of qwen-vl series models (#12458) SijiaYang 2025-11-12 18:08:40 +08:00
  • b40f605fde fix(tcp-port): replace bind_server_socket to get_zmq_socket(Port conflict) (#11961) Jimmy 2025-11-12 18:05:16 +08:00
  • 3cdec20c6b [router] add minmax m2 reasoning parser (#13137) Simo Lin 2025-11-12 18:27:05 +09:00
  • d28caaf60a [router] Support complex assistant and tool messages in /chat/completions (#12860) Danylo Vashchilenko 2025-11-12 00:14:15 -08:00
  • ad8d24c39e Fix re-trigger actor of CI rate limit (#13136) Liangsheng Yin 2025-11-12 16:10:27 +08:00
  • 018123b5b0 [misc] Remove performance and router-benchmark label matching (#13135) Chang Su 2025-11-12 00:04:20 -08:00
  • 4983b7e7aa Fix strict level setting for Kimi K2 tool calls when not explicitly set (#13077) Xinyuan Tong 2025-11-12 00:01:42 -08:00
  • f825137f8a fix: Remove dupulicated kv_events initialization in scheduler (#13132) wxsm 2025-11-12 15:54:47 +08:00
  • ae68158f2e [router] move radix tree to policy crate and addreses some code styles (#13131) Simo Lin 2025-11-12 16:36:04 +09:00
  • a7cc02e36e Fix run suite sanity check (#13133) Ke Bao 2025-11-12 15:29:10 +08:00
  • 8ece99a9dc [CI] Update job dependency and move dpsk v3.2 tests to 8-gpu suite (#12942) Baizhou Zhang 2025-11-11 23:16:47 -08:00
  • 44e391b6e9 [PD] Add custom gpu id to device topo support (#12817) Teng Ma 2025-11-12 15:00:45 +08:00
  • 1a5c313f97 feat(engine): add rid parameter to methods in Engine class (#13095) ishandhanani 2025-11-11 22:37:54 -08:00
  • 7ea5b42d70 [sgl-kernel][5/N]Support Expert Specialization Grouped GEMM (#12666) Qi Yuhang 2025-11-12 13:23:25 +08:00
  • 5ded5e2729 [Feature] Trace: Support http/protobuf span exporter protocol (#12396) zhanghaotong 2025-11-12 13:16:43 +08:00
  • 33d1aeb07f diffusion: refactor task type of models (#13118) Mick 2025-11-12 12:39:34 +08:00
  • dd909a511c Fix gpt oss 4gpu b200 trace links (#12872) alisonshao 2025-11-11 20:31:41 -08:00
  • 60cb716720 [diffusion] log: improve logging while multiprocessing (#12997) Mick 2025-11-12 12:08:37 +08:00
  • 151e13687a Don't fuse wk+weight_proj for nextn (#12863) Trevor Morris 2025-11-11 20:02:52 -08:00
  • 2f9952cdbf diffusion: remove unused workflows folder (#13114) Mick 2025-11-12 11:51:24 +08:00
  • 3e7cc27318 At least tell the user that ngram verify is greedy! (#13039) MayDomine 2025-11-12 11:29:33 +08:00
  • 8f01a12d43 Improve overlap scheduling for better TTFT (#11856) vipwangerxiao 2025-11-12 11:24:04 +08:00
  • 28b8c5792d Upgrade to ROCm 7.0 image (#13105) yctseng0211 2025-11-12 11:14:28 +08:00
  • 7b877ab83d [Router] use call_id instead of id for matching function calls in Responses API for Harmony (#13056) Ziwen Zhao 2025-11-11 17:16:47 -08:00
  • 0d4a418424 [Deepseek V3.2] Fix accuracy bug in the Indexer (#12583) hlu1 2025-11-11 16:15:26 -08:00
  • 2ca25a8aab Revert "fix: display served_model_name in /v1/models" (#13093) Chang Su 2025-11-11 15:58:46 -08:00
  • cc2e36c352 overlap shared + routed expert computation in kimi linear (#12660) b8zhong 2025-11-11 14:52:58 -08:00
  • d8f7816aba [AMD CI] Update nightly docker build CI config. (#13090) Sai Enduri 2025-11-11 14:52:44 -08:00
  • 99e25805f5 [Fix] Fix nan error for large scale ep (#12866) Baizhou Zhang 2025-11-11 14:44:57 -08:00
  • 4a2768a86b Fix spec decoding acc length for dpsk-r1-fp4 tp8 (2nd attempt) (#12915) Qiaolin Yu 2025-11-11 13:30:33 -08:00
  • 9b247f7374 Export runner labels via env var (#13018) Kangyan-Zhou 2025-11-11 13:08:54 -08:00
  • 36d147121f [AMD] Apply AITER_MXFP4_MOE_SF=1 only to gfx950 in aiter build (#13092) Hubert Lu 2025-11-11 12:26:09 -08:00
  • e0e6a6efb2 Fix CPP Radix Cache and add test to CI (#11645) cctry 2025-11-11 11:51:37 -08:00
  • e8114102fa [Fix] Update text_chunks in bench_serving chat completions (#13041) Ziming Huang 2025-11-12 03:41:11 +08:00
  • 14a339fcc8 [Test] Handle streaming chunks with null content in case of stream end. (#10862) vikram singh shekhawat 2025-11-11 23:33:50 +05:30
  • 5f662e786f Revert "[AMD] Add PD test for AMD CI (#11938)" (#13088) Liangsheng Yin 2025-11-12 01:23:55 +08:00
  • 7f5055ed5e Fix cached tokens usage bug (#12814) Frank Minors 2025-11-12 01:02:43 +08:00
  • 63728b1154 [Bug] Login shell error: bash: /root/.cargo/env: No such file or directory (#12941) Johnsonms 2025-11-11 08:52:02 -08:00
  • 5de25f786e [BugFix] Fix prefill memory leak in PD + GDN (#12994) Ziming Huang 2025-11-12 00:48:07 +08:00
  • d52800dbf5 Support file:// scheme in load_video (#13076) Netanel Haber 2025-11-11 18:36:37 +02:00
  • 38a704bccb refine stdout logging codes (#13015) yinghui 2025-11-11 08:16:14 -08:00
  • a06c44f905 fix: display served_model_name in /v1/models (#13063) Edmund Suen 2025-11-12 00:13:58 +08:00
  • 527b7d3f59 [RadixTree] Reduce Stack Push/Pop Overhead for Leaf Nodes, Improve radix_tree Leaf Collection Performance (#12199) PiteXChen 2025-11-11 23:35:39 +08:00
  • f09eee036d Tiny simplify evcition metrics collector (#12983) Liangsheng Yin 2025-11-11 23:23:48 +08:00
  • e38994dd71 Update rope dtype config (#13037) Ke Bao 2025-11-11 22:47:29 +08:00
  • 4a78031a71 [ROCM] Optimized deepseek-r1 model with rmsnorm + fp8 quant fusion (#12689) yctseng0211 2025-11-11 18:59:10 +08:00
  • ea10a9d165 [bug][rocm]fix qr when variable inp (#11609) haoyangli-amd 2025-11-11 17:43:48 +08:00
  • 71aea45c41 [Fix] Add TPOT back to bench_serving (#12976) elvischenv 2025-11-11 17:32:24 +08:00
  • e3b38d7194 [Bug] TypeError: maybe_executor_submit() (#13050) Johnsonms 2025-11-11 00:55:13 -08:00
  • 5c0cadd0c6 Remove duplicate import (#12980) LHXuuu 2025-11-11 16:53:39 +08:00
  • 39b1d048a0 [AMD] Add PD test for AMD CI (#11938) michael-amd 2025-11-11 00:49:44 -08:00
  • 8db7fc4186 disable overlap schedule if mamba radix cache open (#13057) Yi Zhang 2025-11-11 16:35:38 +08:00
  • fe92d4d88e [CI] Auto format code (#13053) Xiaoyu Zhang 2025-11-11 15:30:29 +08:00
  • fc8cda14cf Sglang Tracing: optimize trace_event_batch() (#13036) Feng Su 2025-11-11 15:02:20 +08:00
  • 6a7322ffbc [diffusion] doc: add support_new_models.md (#13043) Mick 2025-11-11 12:49:48 +08:00
  • c751cb38b0 [AMD CI] Update CI Version Logic. (#13029) Sai Enduri 2025-11-10 20:42:00 -08:00
  • 3594815a8b Re-enable Flashinfer TRTLLM GEN MHA and Add Unit Test (#12885) Sam 2025-11-11 12:17:43 +08:00
  • 9caca6a45c [PieceWise CUDA Graph] Support awq/gptq model in piecewise cudagraph (#12518) Xiaoyu Zhang 2025-11-11 11:56:15 +08:00
  • 08c805a85f fix(ci): workflow id in permission rate limit (#13035) yinghui 2025-11-10 19:06:13 -08:00
  • f18ec927f3 fix tuning_fused_moe_triton_sep tool per_channel_quant bug (#13027) Xiaoyu Zhang 2025-11-11 10:33:54 +08:00
  • aea88fa7af [AMD CI] Update docker release workflows docker file name. (#13028) Sai Enduri 2025-11-10 18:28:39 -08:00
  • 2fe4e69fca [router] add postgres databases data connector (#12218) rongfu.leng 2025-11-11 08:51:50 +08:00
  • 012bfc4fdc [9/n] decouple quantization impl from vllm dependency - adjust ci (#12753) Peng Zhang 2025-11-11 06:55:19 +08:00
  • 0493775b06 [router][ci] Quick Improvement to make CI more stable (#12869) Keyang Ru 2025-11-10 13:56:35 -08:00
  • 40b26b456b Simplify the BatchMultimodalOutput in io_struct.py (#12993) Lianmin Zheng 2025-11-10 13:54:56 -08:00
  • 9840bf4f84 [router][ci] Fix maturin build (#13012) Keyang Ru 2025-11-10 12:33:32 -08:00
  • 303cc957e6 chore: bump SGLang version to 0.5.5.post1 (#13000) sglang-bot 2025-11-11 03:53:43 +08:00
  • 661c1c97ad Add pre-suffle weight for new aiter MoE support. (#12908) sogalin 2025-11-11 03:42:55 +08:00
  • c022107f8b Resolve HF download issue and download models before CI run starts for 8-gpu-h200 runners (#12952) Kangyan-Zhou 2025-11-10 11:32:05 -08:00
  • 56c83e0fb3 [CI] Limit the CI trigger frequency of low-privilege actors (#13010) Liangsheng Yin 2025-11-11 03:27:10 +08:00
  • b51d46d092 [AMD CI] Remove SRT docker build. (#11850) Sai Enduri 2025-11-10 11:25:37 -08:00
  • 1086473111 Enhance retract test (page cases, long output cases) (#12781) Liangsheng Yin 2025-11-11 03:03:26 +08:00
  • 665416f6dd Unify memory management across (overlap, non-overlap) x (page>=1) x (spec, non-spec, spec v2) x (retract, finished) (#12224) Liangsheng Yin 2025-11-11 02:56:22 +08:00
  • 838bcb0d93 [misc][ci] Add run-ci after auto-labeler (#13013) Chang Su 2025-11-10 10:41:54 -08:00
  • f1f4c451ab Add process_prefill_chunk back to fix PP event loop (#13009) Liangsheng Yin 2025-11-11 00:51:36 +08:00
  • b0ee99dd03 Super tiny fix typo (#13001) fzyzcjy 2025-11-11 00:47:45 +08:00
  • ddfcb7c8ab minor: fix notebook bug with new model_info fields added for warmup (#13005) Mick 2025-11-11 00:46:12 +08:00
  • 58b12ccb46 Support piecewise cuda graph for deepseek v3 (#12996) Ke Bao 2025-11-10 23:18:03 +08:00
  • 547de8c774 [1 / 2] register weak_ref_tensor in sgl-kernel (#12999) Xiaoyu Zhang 2025-11-10 22:12:59 +08:00
  • 37c40a87a8 chore: bump sgl-kernel version to 0.3.17 (#12966) sglang-bot 2025-11-10 21:50:58 +08:00
  • 1240ac13b8 vlm: fix tiny multimodal cache bug (#12984) Yuhao Yang 2025-11-10 21:30:19 +08:00
  • 5639145fac diffusion: reduce effort of supporting new model (#12982) Mick 2025-11-10 21:20:33 +08:00
  • afee2843d5 feat(metrics): add scheduler and hiradix cache metrics (#10218) (#10225) ShawnKung 2025-11-10 18:14:47 +08:00
  • 6f08488042 fix missing output_token_logprobs when using ngram speculative decoding (#10702) Zhihao Zhang 2025-11-10 18:04:42 +08:00
  • 611a4fd08b [router] bucket policy (#11719) syy-hw 2025-11-10 18:02:53 +08:00
  • 9ea2c686c7 [Auto Sync] Update batch_invariant_ops.py (20251109) (#12916) Lianmin Zheng 2025-11-10 01:51:39 -08:00
  • e2a784ecda [RadixTree] Reduce Syscalls, Optimize Collection Filtering and Align with cpp (#12239) PiteXChen 2025-11-10 17:47:59 +08:00
  • 05559a4a90 Support hidden_dim % 4 == 0 in per_token_quant_fp8 (#12883) Xiaoyu Zhang 2025-11-10 17:13:14 +08:00
  • 7bffc5dc25 [Fix] Add validation for served model name to reserve : for LoRA adapter syntax (#12912) Neelabh Sinha 2025-11-09 23:42:02 -08:00
  • a5e5088dfb Fix errors of page head kernels in sgl-kernel for ROCm (#12604) huangtingwei 2025-11-10 15:15:50 +08:00
  • ac19ce7efb [PP] put pp assert in model runner (#12934) Xuchun Shang 2025-11-10 14:58:19 +08:00
  • 95876d75cb chore: include a minimum image for vlms when warming-up (#9528) Mick 2025-11-10 14:56:59 +08:00
  • a30f190762 chore: bump sgl-kernel version to 0.3.17 (#12931) sglang-bot 2025-11-10 14:54:02 +08:00
  • dc8a5a1ce7 [Refactor / Style] Unify all event loops (except for PP) (#12959) Liangsheng Yin 2025-11-10 14:49:12 +08:00
  • e123648b36 diffusion: fix wan-2.2-TI2V and support sp (#12926) Mick 2025-11-10 14:37:57 +08:00
  • 90401cf7d2 Fix the run-time error when calling fused_rms_mxfp4_quant that change return output number (#12803) kk 2025-11-10 13:48:36 +08:00
  • 9cfe78dd30 clean redundant code in previous PR (#12957) Atream 2025-11-10 13:40:32 +08:00
  • 307e7a6128 diffusion: fix detected file changes rule in CI (#12943) Mick 2025-11-10 13:37:16 +08:00
  • ddd1440d0f Refactor KTransformers heterogeneous compute with unified GPU-quantization backend (#12834) Atream 2025-11-10 13:06:32 +08:00