Commit Graph

  • 4ea4c48bb0 Revert "[ci] fix permission" (#12732) Keyang Ru 2025-11-05 18:26:29 -08:00
  • cf5d27e30a [CPU] Upgrade default PT version to 2.9 (#12611) Zaili Wang 2025-11-06 10:16:35 +08:00
  • 9ec6031d70 [chore]Remove dockerfile from target file of bump kernel version (#12728) Baizhou Zhang 2025-11-05 18:01:19 -08:00
  • a5affb0caf [ci] fix permission (#12729) Keyang Ru 2025-11-05 17:43:48 -08:00
  • 5925d3d719 fix labeler (#12718) Keyang Ru 2025-11-05 17:34:56 -08:00
  • 7ef1964a66 [router] add basic ci tests for gpt-oss model support (#12651) Keyang Ru 2025-11-05 15:45:11 -08:00
  • b0476a0642 [router][quick fix] Add minimal option for reasoning effort in spec (#12711) Keyang Ru 2025-11-05 15:43:19 -08:00
  • ffba61a100 [router][grpc] Make harmony parser checks recipient first before channel (#12713) Chang Su 2025-11-05 15:33:52 -08:00
  • 3c219eb0e4 [misc] Change sync-labels to false (#12714) Chang Su 2025-11-05 15:33:33 -08:00
  • 1ffdcdc430 Revert "Commented out b200 tests due to runner shortage (#12609)" (#12712) Kangyan-Zhou 2025-11-05 15:09:49 -08:00
  • c7d57d5bb3 Fix CI and style (#12658) Lianmin Zheng 2025-11-05 15:08:15 -08:00
  • 80802c4cc6 [router][ci] speed up python binding to 1.5 min (#12673) Keyang Ru 2025-11-05 14:52:33 -08:00
  • 83b104ee8c [misc] Add labeler for automatic labeling (#12710) Chang Su 2025-11-05 14:44:43 -08:00
  • 141278048e [NVIDIA] Fix unit test of MoE and add it to nightly ci (#12709) Kaixi Hou 2025-11-05 14:33:18 -08:00
  • 82f39dc11d Add mm_fp4 trtllm backend (#12406) Shu Wang 2025-11-05 14:31:46 -08:00
  • 627bac649c Support Expert Deferral Mechanism in KTransformers (#12586) Atream 2025-11-06 05:41:52 +08:00
  • 3651cfbf62 [router] fix: validate HTTP status codes in health check (#12631) wyx 2025-11-06 05:37:17 +08:00
  • c8547ecddd Enable Aiter Attention for VL model (#12699) Morpheus Guo 2025-11-06 05:01:23 +08:00
  • 7bc1dae095 WIP: initial multimodal-gen support (#12484) Mick 2025-11-06 04:28:52 +08:00
  • 4fe53e5888 [router][grpc] Support streaming parsing with Tool Choice in chat completions API (#12677) Chang Su 2025-11-05 12:12:21 -08:00
  • fb2e816e83 Fix server args for gpt oss so users can override the moe runner backend (#12696) Lianmin Zheng 2025-11-05 11:36:59 -08:00
  • 7c45b8b4bb [CI] Fix qwen3-vl lora nightly ci (#12708) Baizhou Zhang 2025-11-05 11:00:13 -08:00
  • ba5b68236f Commented out b200 tests due to runner shortage (#12609) Kangyan-Zhou 2025-11-05 10:47:43 -08:00
  • 508d2f7aa2 add Kimi k2 reasoning parser (#12702) bigmoyan 2025-11-06 00:37:54 +08:00
  • a889c85459 [Grammar Fix] GLM-4-MOE self.first_k_dense_replace is undefined. (#12455) Yuxuan Zhang 2025-11-06 00:03:45 +08:00
  • 4d84f886e7 Refactor --debug-tensor-dump-layers to list (#12691) Yuhong Guo 2025-11-05 19:30:01 +08:00
  • dc4f541823 fix trtllm_mla attention backend when disabling cuda graph. (#12687) yinghui 2025-11-05 01:35:02 -08:00
  • 0648eb482d [Profiler] Add SGLANG_PROFILE_RECORD_SHAPES for recording shapes when profiling (#11641) zejunchen-zejun 2025-11-05 15:41:46 +08:00
  • b88fab3111 fix: add seed bench_serving to cache key, remove redundant function definition. (#12680) yinghui 2025-11-04 23:39:11 -08:00
  • 3694266051 Expand and update test coverage for AMD CI (#10044) Hubert Lu 2025-11-04 22:15:13 -08:00
  • 9f5e701879 [router][grpc] Implement tool_choice support for Responses API (#12668) Chang Su 2025-11-04 21:43:42 -08:00
  • cbf23dbbfa [Feature] add --lora-request-distribution arg to bench_serving.py and support skewed and distinct workloads (#12175) Glen Liu 2025-11-05 00:41:40 -05:00
  • 6dade6c3b5 Fix VLLM dependency test (#12670) Kangyan-Zhou 2025-11-04 20:49:58 -08:00
  • b419e20c5b [Dockerfile] Speed up docker image building (#8784) Yingchun Lai 2025-11-05 12:26:18 +08:00
  • 48641435d6 fix typo of args description in sglang.profiler (#12486) ai-easy-cpu 2025-11-05 12:15:13 +08:00
  • 44b1b394a4 [PD-Disagg] Check finish after pop tranferred (#12638) Liangsheng Yin 2025-11-05 11:18:09 +08:00
  • 0711d1509b [NVIDIA] Fix cutedsl backend of MoE (#12353) Kaixi Hou 2025-11-04 18:54:55 -08:00
  • 09938e1f82 chore: bump SGLang version to 0.5.4.post3 (#12639) sglang-bot 2025-11-05 10:32:11 +08:00
  • 2340798353 Register allgather/reducescatter buffers with symm memory (#12572) Nicolas Castet 2025-11-04 19:11:36 -06:00
  • 1357ab025a [router][grpc] Emit OutputItemDone event and store output item array (#12656) Chang Su 2025-11-04 17:01:03 -08:00
  • 44da737770 [fix] Handle escaped characters in GLM tool call parser to prevent double serialization (#12456) soaringk 2025-11-05 08:48:14 +08:00
  • fb9582c4e1 Add multi-GPU configurations to nightly-test.yml (#12585) alisonshao 2025-11-04 16:46:30 -08:00
  • d22d044734 Revert "Enable memory saver for hybrid model" (#12648) Baizhou Zhang 2025-11-04 16:22:06 -08:00
  • 887742a1e7 [router][grpc] Fix index issues in reasoning content and missing streaming events (#12650) Chang Su 2025-11-04 15:38:43 -08:00
  • 34f7564df0 [NVIDIA] Fix wrong symmetric sizes for fp4 cases (#12640) Kaixi Hou 2025-11-04 14:19:37 -08:00
  • 1cfbbc42d8 [Bug] Fix NSA Backend KV-Buffer Shape Mismatch in DeepSeek-V3.2 (#12645) Johnsonms 2025-11-04 13:57:32 -08:00
  • 55dfb539cf [Auto Sync] Update scheduler_metrics_mixin.py, collector.py (20251104) (#12647) Lianmin Zheng 2025-11-04 13:56:14 -08:00
  • 42889acbd0 [hotfix] Fix deepep w4a8 bug (#12642) Baizhou Zhang 2025-11-04 13:55:59 -08:00
  • 211f4070e5 fix: Lazy import mooncake-ep to fix extra gpu contexts being created (#12641) Trevor Morris 2025-11-04 12:28:36 -08:00
  • befa41a152 Fix output_ids inconsistency (#12628) Liangsheng Yin 2025-11-05 01:43:08 +08:00
  • 30b26ee9d0 Add io struct naming check back (#12634) Liangsheng Yin 2025-11-05 01:15:01 +08:00
  • aa797d013d [Test] Merge all constrained decoding tests. (#12633) Liangsheng Yin 2025-11-05 00:43:06 +08:00
  • 7cee07a067 Fix skip layer in get_quant_method (#12632) Ke Bao 2025-11-04 23:27:46 +08:00
  • bb517fe393 [HotFix] Disable torch dynamo for mrope_triton kernel (#12593) Yuan Luo 2025-11-04 23:26:56 +08:00
  • 0e82fd3df4 [router][grpc] Fix model validation, tool call check, streaming logic and misc in responses (#12616) Chang Su 2025-11-04 02:21:58 -08:00
  • b7d7041190 Add sanity checks when a test file is not added to CI (reland) (#12594) fzyzcjy 2025-11-04 18:04:26 +08:00
  • ff0b64e1e6 Ensure GPU work is finished when release memory occupation call is finished (#12592) fzyzcjy 2025-11-04 18:01:27 +08:00
  • d84790db39 Support aggregating engine metrics in sgl-router (#11456) fzyzcjy 2025-11-04 17:59:50 +08:00
  • 0678beaaee [sepc-v2] Fix imcompatibility with constrained decoding (#12615) Liangsheng Yin 2025-11-04 17:27:31 +08:00
  • c2d4716da0 chore: bump mooncake version to 0.3.7.post2 (#12599) Shangming Cai 2025-11-04 17:09:31 +08:00
  • c14cc47e39 [Deterministic] Optimize bmm_batch_invariant op (#12522) Minglei Zhu 2025-11-04 00:33:31 -08:00
  • dbcf85b7f0 Add --speculative-moe-runner-backend server arg (#10183) Trevor Morris 2025-11-04 00:20:56 -08:00
  • 83804bc626 [router][grpc] Restructure modules and code clean up (#12598) Chang Su 2025-11-03 23:59:27 -08:00
  • d5fa019c36 feat: limit peak memory usage when computing logprobs (#6318) Zhao Chen 2025-11-04 15:53:20 +08:00
  • fef3a6b63b Restore torch defaults between sgl-kernel tests (#11131) Ben Barsdell 2025-11-04 18:51:23 +11:00
  • 173e0f704f Enable memory saver for hybrid model (#11974) Junrong Lin 2025-11-04 14:55:26 +08:00
  • f600866a44 Improve the metrics for PD (#12580) Lianmin Zheng 2025-11-03 22:10:57 -08:00
  • 93be7e863e fix: respect --ignore-eos in PD case for benchmarking (#12597) ishandhanani 2025-11-03 21:44:14 -08:00
  • 60b0754cc9 Tiny fix ExpertDistributionReq error (#11760) fzyzcjy 2025-11-04 13:39:25 +08:00
  • 0b24af4d79 test: support return logprobs in bench_offline_throughput test (#12462) Zhao Chen 2025-11-04 13:38:48 +08:00
  • a209fb05c1 [Qwen3 VL] Add LoRA support for Qwen 3 VL (#12165) Jonah Bernard 2025-11-03 23:32:54 -05:00
  • 48d6bea1ea [GDN/SWA] mamba and swa radix cache edge case fix (#12111) Hanming Lu 2025-11-03 19:03:37 -08:00
  • 1689c0e35f [Doc] fix miss index for production request trace (#12547) Teng Ma 2025-11-04 09:57:09 +08:00
  • 193fbb0bce Super tiny add UT for copy_to_gpu_no_ce (#12270) fzyzcjy 2025-11-04 09:40:51 +08:00
  • e607850fcf Enable mixed type LayerNorm kernel for NSA indexer (#12044) akhilg-nv 2025-11-03 16:50:41 -08:00
  • 15efbcb4e7 [chore] Fix update_kernel_whl_index script for multiple cuda version (#12519) Baizhou Zhang 2025-11-03 16:34:14 -08:00
  • 243c064df2 Remove the dependency of nccl.h in symmetric memory (#12571) Lianmin Zheng 2025-11-03 16:11:00 -08:00
  • 0b41a293fa [router][grpc] Consolidate error messages build in error.rs (#12301) Chang Su 2025-11-03 14:41:45 -08:00
  • d31d48b341 update usage of trtllm_fp8_per_tensor_scale_moe (#12569) b8zhong 2025-11-03 14:25:32 -08:00
  • 8834260739 Super tiny dump server info such as args in bench for post analysis (#12550) fzyzcjy 2025-11-04 06:24:08 +08:00
  • fd7a72d62d Super tiny allow profile activities in bench_serving (#12549) fzyzcjy 2025-11-04 06:23:18 +08:00
  • 21a8fa16ea tiny optimize for bench serving (#12553) Yi Zhang 2025-11-04 06:13:18 +08:00
  • 7a21d8b276 Reduce the overhead of nccl symmetric memory (#12524) Lianmin Zheng 2025-11-03 11:56:27 -08:00
  • d36639eec7 [ROCm] Update Mooncake to v0.3.7.post1 and add -DUSE_HIP=ON to rocm.Dockerfile (#12560) R0CKSTAR 2025-11-04 03:33:30 +08:00
  • 6ef23b9833 [Test] Add parameters to SRTRunner (#12227) Jonah Bernard 2025-11-03 14:20:56 -05:00
  • 385599cb04 Fix error when calling quantization (#12548) fzyzcjy 2025-11-04 02:17:43 +08:00
  • 952fbe47cb fix: fix the bug which leads qwen2_5_vl to crash with mixed_chunk (#11330) Yueyang Pan 2025-11-03 18:26:03 +01:00
  • edb2569356 [hot-fix] Fix broken CI (#12564) Liangsheng Yin 2025-11-04 00:03:25 +08:00
  • 3529c061bb [spec v2] Fix output repetition by speculative sampling error (#12561) Liangsheng Yin 2025-11-03 23:00:17 +08:00
  • ffb32a8548 Conditionally recapture cuda graph after model weight update from disk (#12060) harrisonlimh 2025-11-03 05:51:27 -08:00
  • 14d8064803 fix: Fix KTransformers hybrid inference with int8 quantization and format (#12536) Atream 2025-11-03 20:59:39 +08:00
  • ab8b83f71d chore: upgrade mooncake 0.3.7.post1 (#12541) Shangming Cai 2025-11-03 16:34:32 +08:00
  • de0b10cf5c fix: move dummy format loader check before quantization checks (#12532) yinghui 2025-11-02 23:41:30 -08:00
  • 6e29446e45 [hotfix] Remove flashinfer-jit-cache from pyproject (#12530) Baizhou Zhang 2025-11-02 22:11:05 -08:00
  • 0c3543d7d5 chore: upgrade flashinfer 0.5.0 (#12523) Yineng Zhang 2025-11-02 20:54:12 -08:00
  • 6a3b9fd00f Update setup_github_runner.md Kangyan-Zhou 2025-11-02 20:44:09 -08:00
  • 65f1d065c5 [Bug] Fix Intern-S1 model accuracy and support /generate interface with input_ids (#12367) Haian Huang(深度眸) 2025-11-03 12:22:33 +08:00
  • 9434a0e50f [Refact] Remove hardcoded KV cache dimension in MLATokenToKVPool (#12502) Johnsonms 2025-11-02 19:49:53 -08:00
  • 20315697f4 move all get_stream in sgl_kernel to c++ to reduce the launch overhead (#12521) Lianmin Zheng 2025-11-02 13:15:05 -08:00
  • c9db79117f Super tiny fix naming in bench serving scripts (#12515) fzyzcjy 2025-11-03 04:43:10 +08:00