Commit Graph

  • b12b40dea4 Add cuda_graph_forward_passes_total and num_retracted_reqs_total (#15189) fzyzcjy 2025-12-17 21:19:42 +08:00
  • 6c4bf8a0be [diffusion] profiling: enhance trace export with gzip and integrity check (#15326) Xiaoyu Zhang 2025-12-17 20:44:55 +08:00
  • 533851fbcb [diffusion] ci: add flux2 tp2 test into ci to avoid breaking tensor parallel (#15237) Xiaoyu Zhang 2025-12-17 20:44:21 +08:00
  • 0071fe9c40 [bug fix][pp] fix weight load for qwen2.5-vl (#15138) Xuchun Shang 2025-12-17 18:33:50 +08:00
  • ffa7e03506 [Piecewise CUDA Graph] Support INT8 (#14918) b8zhong 2025-12-17 02:20:57 -08:00
  • 712f44ee2b fix qwenvl compressed tensors quantization weight loader (#11914) LHXuuu 2025-12-17 18:01:43 +08:00
  • 8c34e18140 Fix the accuracy issue when running mxfp4 dsv3 model and enable ep (#15304) kk 2025-12-17 17:53:22 +08:00
  • e9abb52576 [bug fix][pp] fix qwen3 model load (#15223) Xuchun Shang 2025-12-17 17:26:22 +08:00
  • 888594333e Fix gpu-fault when running mtp in eager mode (#15233) kk 2025-12-17 17:25:06 +08:00
  • feb8e30b9d [Hotfix] Fix required enable_mamba_track argument for Flashinfer autotune path (#15314) elvischenv 2025-12-17 17:03:24 +08:00
  • 45a959d3e9 [PP] Add pp support for Qwen3-VL (#12333) Xuchun Shang 2025-12-17 16:03:58 +08:00
  • cdce516331 [diffusion] api: add sampling parameters and model info endpoint to OpenAI API (#15071) WenhaoZhang 2025-12-17 15:33:18 +08:00
  • 79ab57bd7a Revert "direct register custom op for mm_fp4 (#13699)" (#15284) b8zhong 2025-12-16 23:09:44 -08:00
  • 4b8901ac0f Update FP4 GEMM Benchmark (#14449) b8zhong 2025-12-16 23:04:56 -08:00
  • 435d1c83c1 [Perf] Enable Flashinfer autotune by default (#14357) elvischenv 2025-12-17 15:01:39 +08:00
  • 2bdbaef18e [DeepSeekV3.2] Add pure TP+MTP test (#15088) Ashton Chew 2025-12-16 21:48:12 -08:00
  • 31d48d7f6f Add Ollama-compatible API endpoints + Smart Router (#14376) Alison Shao 2025-12-16 20:43:38 -08:00
  • 0129c911e0 fix(function_call): fallback to decode when batch decode options differ (#15155) luqitao 2025-12-17 12:21:08 +08:00
  • 03f9eb2564 [BugFix] fix gptq_marlin_gemm has no parameter called b_bias (#13571) ehuaa 2025-12-17 10:03:55 +08:00
  • 7ec678eb09 [Test] Update LoRA eviction policy tests to match current behavior (#15283) Alison Shao 2025-12-16 17:55:15 -08:00
  • 9d64a7b24f Minor style fixes to the scheduler.py (#15218) Lianmin Zheng 2025-12-16 17:09:44 -08:00
  • 46ad4b986d fix: moving decorator to header (#15297) Douglas Yang 2025-12-16 17:05:45 -08:00
  • 71cb90378b [NPU] fix for NPU memory settings logic (#15258) Even Zhou 2025-12-17 09:04:22 +08:00
  • c8c64876a7 [NVIDIA] Fixes for NVFP4 all-gather with spec decoding (#15280) Trevor Morris 2025-12-17 00:58:39 +00:00
  • 0861dca81f Revert "[misc] Upgrade cutedsl to 4.3.1 (#14857)" (#15293) Yineng Zhang 2025-12-16 16:31:32 -08:00
  • da58df6b3d fix: skipping TestEPDDisaggregationOneEncoder test (#15292) Douglas Yang 2025-12-16 16:31:20 -08:00
  • 49237e26fb Fix test_pp_single_node.py estimated time from 800s to 500s (#15291) Alison Shao 2025-12-16 16:27:09 -08:00
  • d92c1f8cbd Fix lora doc (#15282) Baizhou Zhang 2025-12-16 15:53:38 -08:00
  • 930705863f [sgl-kernel] Update flashmla to include fp8 sparse_mla optimizations (#15242) hlu1 2025-12-16 15:12:19 -08:00
  • ccc8f3b266 support non disturbing remote instance weight loader v2 (#14997) amysaq2023 2025-12-17 06:39:56 +08:00
  • a4c762811a Remove incorrect BlockRemoved event emission during node splits (#14934) Neal Vaidya 2025-12-16 14:11:24 -08:00
  • 28a19e494b Fix lint (#15281) Baizhou Zhang 2025-12-16 13:36:01 -08:00
  • 0261c4aff7 [misc] Upgrade cutedsl to 4.3.1 (#14857) Baizhou Zhang 2025-12-16 12:11:56 -08:00
  • 8ac350f335 [AMD] Support fused_rms_mxfp4_quant in the prefill stage for DeepSeek-R1-MXFP4 (#14975) jacky.cheng 2025-12-17 04:03:58 +08:00
  • 99401e7b1a Fix accuracy issue when using a16w16 mla_decode_fwd (#14936) kk 2025-12-17 01:55:18 +08:00
  • 6682475124 [CI] Improve flaky 4 GPU test success rate (#15234) Shangming Cai 2025-12-17 00:02:26 +08:00
  • f95729b06f [diffusion] doc: update profiling.md (#15270) Mick 2025-12-16 23:48:32 +08:00
  • 9f4ed93dd8 [diffusion] multi-platform: use current_platform.device_type to replace hard-coded cuda device (#15232) R0CKSTAR 2025-12-16 22:17:33 +08:00
  • ecb401ed42 Enhance runtime memory check in CI (#15192) Liangsheng Yin 2025-12-16 21:40:38 +08:00
  • b399e3ac4f Support piecewise cuda graph for fused marlin moe (#15100) Ke Bao 2025-12-16 20:05:32 +08:00
  • e27635a02d [CPU] Add 4D input support for ROPE in sgl-kernel (#9337) blzheng 2025-12-16 17:27:39 +08:00
  • 272c5fe43e Increase timeout for TestDeepseekV3MTP for potential DeepGEMM cold start (#15239) Kangyan-Zhou 2025-12-16 00:26:10 -08:00
  • 3c8dc448b2 fix: removing latest-sglang=1 (#15220) Douglas Yang 2025-12-16 00:14:48 -08:00
  • 5e96beb3e5 Adding tool calling and reasoning parser support for Intern-S1 (#14866) Kenny Yao 2025-12-16 03:01:58 -05:00
  • 9327482baa [bugfix][quark] Fixed an issue where per_token could not be properly recognized when the token count was 1. (#14415) haoyangli-amd 2025-12-16 14:54:31 +08:00
  • 36fcf71fff [Qwen3-next] Add PD disaggregation support for mamba with extra_buffer (#15180) Shangming Cai 2025-12-16 14:36:00 +08:00
  • 6292d97135 [diffusion] fix: fix pack qkv opt break tensor parallel (#15225) Xiaoyu Zhang 2025-12-16 14:33:49 +08:00
  • c843419562 Remove duplicate bs=1 in nightly benchmark (#15162) Baizhou Zhang 2025-12-15 22:09:22 -08:00
  • 3e4d431a44 [Feature] Add AIME25 dataset support for SGLang simple_eval (#14990) ゆり 2025-12-16 14:59:40 +09:00
  • 538e733e08 [AMD CI] Fix typo. (#15229) Sai Enduri 2025-12-15 21:47:01 -08:00
  • 22587bc0b4 [BugFix] Fix CPU inference failure (#15231) Yibo Cai 2025-12-16 13:32:41 +08:00
  • 02d24244e4 [AMD CI] Temporarily disable 2 gpu accuracy test. (#15204) Sai Enduri 2025-12-15 20:57:26 -08:00
  • 1da5cd638f [Bugfix][Tool Call] Add null system prompt to support tool system prompt (#15092) Muqi Li 2025-12-16 11:40:10 +08:00
  • a9a2cdd8ec Add EPD disaggregation doc (#15224) Tianyu Guo 2025-12-16 11:22:05 +08:00
  • 4733fcff1f [Feature] npu support enable_torch_compile for torchair backend (#13410) XDaoHong 2025-12-16 09:23:51 +08:00
  • 3ffa260474 chore: update CI_PERMISSIONS (#15212) Yineng Zhang 2025-12-15 15:09:05 -08:00
  • 61f362c694 feature: create docker image from pr branch (#15185) Douglas Yang 2025-12-15 10:53:50 -08:00
  • e7157c9b77 Add cache for flashinfer installation (#15153) Kangyan-Zhou 2025-12-15 10:37:34 -08:00
  • 30da2f0598 [NPU][eagle3] support qwen eagle3 on NPU (#14820) Liwansi 2025-12-16 02:25:13 +08:00
  • 4901693110 [diffusion] perf: support FFN pack gate and up proj for Z-Image(#15201) Xiaoyu Zhang 2025-12-16 01:18:47 +08:00
  • 3d484be547 fix(attention): Prevent trtllm_mha auto-selection with eagle3 speculative decoding (#15127) ratish 2025-12-15 20:42:26 +04:00
  • c0d94440b7 [diffusion] perf: support pack qkv for Z-Image (#15191) Xiaoyu Zhang 2025-12-16 00:22:24 +08:00
  • 1dedb63860 [diffusion] chore: minor code cleanups (#15190) Mick 2025-12-15 23:57:02 +08:00
  • 7bc8b1532e [diffusion] fix: fix AttributeError in _build_parallelism_config when accessing tp_group.device_group (#15196) Xiaoyu Zhang 2025-12-15 23:41:06 +08:00
  • 9003a4369d Add missing assertion in NemotronH path (#15193) roikoren755 2025-12-15 16:17:21 +02:00
  • 3518b33178 [model-gateway] Remove legacy RouterMetrics and Rename SmgMetrics to Metrics and smg_labels to metrics_labels (#15160) Simo Lin 2025-12-15 04:57:42 -08:00
  • b098b1ae24 [diffusion] fix: fix video model sp when resolution is not specified (#15047) Mick 2025-12-15 20:25:43 +08:00
  • abd3e048f2 [diffusion] fix: fix pytorch non-writable array warning (#15017) Lancer 2025-12-15 20:18:54 +08:00
  • 92c29d43ac [diffusion] fix: cache dit with parallel (#15163) Xiaoyu Zhang 2025-12-15 19:15:51 +08:00
  • bf6438142a chore: change npu pr-test a2 runner (#15152) Goalina 2025-12-15 19:03:19 +08:00
  • f03bfa4ce3 [Feature] Fuse mrope all in 1 kernel (#14906) DarkSharpness 2025-12-15 18:50:55 +08:00
  • 89ad390843 Fix num running requests (load) wrong cleared for ongoing requests (#15116) fzyzcjy 2025-12-15 17:40:04 +08:00
  • 2ea844ec81 Fused two elementwise kernels for k_nope and k_pe concat (#14862) kk 2025-12-15 17:33:33 +08:00
  • 1e2d753804 fix: adding date and fixing release name issue (#15174) Douglas Yang 2025-12-15 01:11:03 -08:00
  • d16ff357db [CPU] Add Gemma3RMSNorm kernel in sgl-kernel and add ut (#9324) blzheng 2025-12-15 16:24:02 +08:00
  • af49e30242 feature: PR wheel (#15170) Douglas Yang 2025-12-15 00:22:47 -08:00
  • 01b955ac3d [diffusion] model: support mutli-image input and qwen-image-edit-2509 (#15005) Yuhao Yang 2025-12-15 16:17:10 +08:00
  • 16e6bc20b0 fix CompressedTensorsW8A8Int8 min_capability (#13914) mmdbhs 2025-12-15 15:50:39 +08:00
  • 372507649d Tiny improve summary text in bench_one_batch_server.py (#15158) Liangsheng Yin 2025-12-15 14:59:11 +08:00
  • 7b9156c773 [model-gateway] add mcp and discovery metrics (#15156) Simo Lin 2025-12-14 22:54:53 -08:00
  • 21cfebac65 fix: move ci-bot (#15154) Douglas Yang 2025-12-14 22:47:19 -08:00
  • 702426b06a Fix import warnings (#15144) Lianmin Zheng 2025-12-14 21:25:18 -08:00
  • 9e9a61691e ci: adding errors to Github summary (#14778) Douglas Yang 2025-12-14 21:08:16 -08:00
  • bd9c3a47d6 [model-gateway] Add streaming metrics for harmony gRPC router (#15147) Simo Lin 2025-12-14 20:54:43 -08:00
  • 1e641ee4a7 [model-gateway] upgrade axum and axum server (#15146) Simo Lin 2025-12-14 20:54:01 -08:00
  • fb96669ff9 [model-gateway] Add Layer 3 worker metrics (smg_worker_*) (#15130) Simo Lin 2025-12-14 20:25:46 -08:00
  • 3912ee4991 [VLM] feat: support chunked vit attention (#14907) Yuan Luo 2025-12-15 12:11:02 +08:00
  • 1ab9b8e0a3 Enable TRT AllReduce Fusion by default (#14764) b8zhong 2025-12-14 20:01:49 -08:00
  • e61dabf5e4 [Qwen3-next] support mamba radix cache for overlap scheduler (#14792) Hanming Lu 2025-12-14 18:54:16 -08:00
  • 36e7c8c59f docs: update usage (#15142) Yineng Zhang 2025-12-14 18:42:21 -08:00
  • 037c3982af Fix H200 CI by commenting out Warmup Weights and JIT Compilation (#15139) Kangyan-Zhou 2025-12-14 18:04:28 -08:00
  • 62b3fdae43 Fix cache aware wrong routing caused by incorrect load tracking (#15101) fzyzcjy 2025-12-15 10:03:54 +08:00
  • 1cd0c3bfb7 Add cyb70289 to CI permissions (#14938) Yibo Cai 2025-12-15 09:51:30 +08:00
  • 2c899431f8 fix: adjusting frequency for ci failure monitor (#15134) Douglas Yang 2025-12-14 17:49:57 -08:00
  • 4513f549ee [diffusion] fix: fix default resolution 720p width from 1080 to 1280 (#15058) Xiaoyu Zhang 2025-12-15 09:16:47 +08:00
  • c96903074c [NSA] Fix NSA backend assertion error when running DeepSeek-V3.2 PP with radix-cache (#15086) YAMY 2025-12-14 17:13:18 -08:00
  • 4449c17011 [model-gateway] fix circuit breaker metrics (#15099) fzyzcjy 2025-12-15 07:43:04 +08:00
  • 5ca962ce7f [model-gateway] extract circuit breaker state struct (#15098) fzyzcjy 2025-12-15 07:40:23 +08:00
  • bab20a849e [model-gateway] Parallelize metrics requests (#14953) Praneth Paruchuri 2025-12-15 05:08:45 +05:30
  • 0e4108ba29 feat(gateway): Add server-side TLS support (#15052) ratish 2025-12-15 03:37:55 +04:00