Commit Graph

  • 67e6f1438d Consolidate similar tests to reduce duplication (#12871) alisonshao 2025-11-14 22:29:44 -08:00
  • 34851471b2 Add more statistics for spec decoding (#13317) Zilin Zhu 2025-11-15 14:20:28 +08:00
  • c2083116e3 [Diffusion] switch to local calculate_dimensions (#13294) Adarsh Shirawalmath 2025-11-15 11:25:13 +05:30
  • 10285ec204 [Misc]Add date to cu13 dev image tag (#13316) Baizhou Zhang 2025-11-14 19:42:50 -08:00
  • 9b3fc186bc Fix syntax errors in cpp_radix_tree (#13315) Yiming 2025-11-15 11:35:43 +08:00
  • 172c71a23d [model-gateway] smg release 0.2.3 (#13312) Simo Lin 2025-11-14 19:04:20 -08:00
  • af373636da CI: add server performance test for SGLang diffusion (#13091) Adarsh Shirawalmath 2025-11-15 07:41:19 +05:30
  • a5be6ef98e [Deterministic] Support Qwen3-Next model deterministic inference (#13100) Minglei Zhu 2025-11-14 18:01:18 -08:00
  • 8f4e18a294 Add missing model for 2-gpu-runner in nightly tests (#13311) Kangyan-Zhou 2025-11-14 17:34:43 -08:00
  • e9681444bd Add missing model in model validate list (#13310) Kangyan-Zhou 2025-11-14 17:29:21 -08:00
  • eae59b337e Piecewise Cuda Graph Support for gpt-oss model (#13045) Yuwei An 2025-11-14 17:28:00 -08:00
  • 14dc052382 [router]Replace requests lib with openai in e2e_response_api (#13293) Xinyue Zhang 2025-11-14 16:56:40 -08:00
  • b223669136 Remove nightly b200 tests and revert a change for test file (#13305) Kangyan-Zhou 2025-11-14 16:30:22 -08:00
  • dcc47a56c9 Implement nightly test workflow naming conventions (#13170) alisonshao 2025-11-14 16:14:53 -08:00
  • 6448b4cd2c Fix NSA indexer nightly test failed issues (#13298) Johnsonms 2025-11-14 13:14:57 -08:00
  • 5ae0ac4244 [NVIDIA] Fix use case of SGLANG_ENABLE_FLASHINFER_GEMM (#13274) Kaixi Hou 2025-11-14 12:51:11 -08:00
  • 22f641ab4f [minor] remove debug code in python/sglang/srt/compilation/weak_ref_tensor_jit.py (#13235) Lianmin Zheng 2025-11-14 12:22:24 -08:00
  • 0997c78d2c Support FP8 Per Token Quant Piecewise (#13272) Stefan He 2025-11-14 11:40:12 -08:00
  • fd3be107bb [Doc] Add item for repetition punishment (#13260) Zesen SenmiaoORZ 2025-11-14 14:15:56 -05:00
  • 665f43bdd8 model: support teleflm (#10573) Praneth Paruchuri 2025-11-15 00:44:49 +05:30
  • a53f2d6c12 Support orion (#10665) Praneth Paruchuri 2025-11-15 00:38:32 +05:30
  • a7002e614b [Deepseek V3.2] Clean up MTP (#13236) hlu1 2025-11-14 11:01:37 -08:00
  • 84e151ac7f LLama4 Attention: Update assertion msg (#12777) Kshiteej K 2025-11-14 19:52:50 +01:00
  • 875a25ddef refactor: remove duplicate function _get_bootstrap_info_from_server (#13277) Yingchun Lai 2025-11-15 02:48:17 +08:00
  • fc55b45e5f Optimized prefill cache allocation for NPU (#13288) Vitaly Tuzov 2025-11-14 21:40:17 +03:00
  • 0050ff254f [BugFix] fix bench_serving error when multimodal image is testing (#13254) Ling Zhang 2025-11-14 22:46:19 +08:00
  • af9f71f9c5 Add script to create a model with fewer layers for debugging (#13284) fzyzcjy 2025-11-14 22:13:27 +08:00
  • 15264232ee Super tiny fix CI (#13283) fzyzcjy 2025-11-14 21:45:10 +08:00
  • f8d3d80f63 chore: bump flashinfer v0.5.2 (#13242) Yineng Zhang 2025-11-14 02:47:09 -08:00
  • 5027739f2c [CPU] Use covt_e4m3_bf16 to optim BF16 to FP8 convert (#12191) wangyxbh 2025-11-14 01:36:51 -08:00
  • 3701f34dab Tiny add utility to parse server logs (#12605) fzyzcjy 2025-11-14 17:34:22 +08:00
  • 821fb060c3 Enhance dumper comparator with tensor unifier and location finder (#12623) fzyzcjy 2025-11-14 17:34:08 +08:00
  • ace27c0c01 Tiny enhance dumper with ctx and enable flags (#12622) fzyzcjy 2025-11-14 17:33:49 +08:00
  • ed1d18d472 Tiny fix update version logic location (#12620) fzyzcjy 2025-11-14 17:33:36 +08:00
  • e7b57b0d04 [BugFix] weight load bug when checkpoint expert.gate and exepert.up_proj are not fused (#13113) Morpheus Guo 2025-11-14 16:35:44 +08:00
  • 49141df94a Extend lint test to test/ directory (#13247) Kangyan-Zhou 2025-11-14 00:01:48 -08:00
  • 7b79cc4fe2 ci: speed up b200 ci (#13237) b8zhong 2025-11-13 23:57:33 -08:00
  • 385ff0e56f [model-gateway] change mg labeler from router to model-gateway (#13265) Simo Lin 2025-11-14 16:29:09 +09:00
  • fc5da1e80b [Tool Call] Steamline function arguments when tool_choice="auto" for deepseekv31_detector (#11589) Muqi Li 2025-11-14 15:06:02 +08:00
  • 04848ba7cb Add 3 models to 2 gpu runner in model downloading from nightly tests (#13261) Kangyan-Zhou 2025-11-13 22:59:18 -08:00
  • 9bc6a9adbe Update model weight validation logic to handle special weight file naming (#13256) Kangyan-Zhou 2025-11-13 22:07:49 -08:00
  • 7cdaedb8fb Remove glm41v from CI to speed up CI (#13257) Binyao Jiang 2025-11-13 21:14:04 -08:00
  • 1f134f850a fix outdated router doc (#13255) fzyzcjy 2025-11-14 13:00:42 +08:00
  • 5f72d36d40 Fix nightly tests to fail properly when any job fails (#13096) alisonshao 2025-11-13 19:10:35 -08:00
  • 7a2254b260 [Auto Sync] Update backend.py, forward_batch_info.py, piece... (20251113) (#13221) Lianmin Zheng 2025-11-13 18:54:11 -08:00
  • b5904999c0 docker: Fix apt-add-repository (#13213) Kshiteej K 2025-11-14 03:41:29 +01:00
  • 922525ee6d fix: fix serve command without diffusion dependency (#13246) Mick 2025-11-14 10:24:18 +08:00
  • ee3e337c62 [feat] make warmup timeout configurable through SGLANG_WARMUP_TIMEOUT (#13243) billishyahao 2025-11-14 10:14:37 +08:00
  • e523e2167a remove deprecated tile_tokens_dim (#13186) b8zhong 2025-11-13 17:09:27 -08:00
  • 2ce2377740 [AMD] Fix AITER_MXFP4_MOE_SF setting for gfx950 (#13239) Hubert Lu 2025-11-13 17:03:41 -08:00
  • 19f6a33ce9 [Auto Sync] Update pynccl_wrapper.py, environ.py, registry.... (20251111) (#13097) Lianmin Zheng 2025-11-13 15:58:47 -08:00
  • e9c0c55833 [sgl-kernel] clean up fa fetch in CMakeLists.txt (#12392) Fan Yin 2025-11-14 07:42:09 +08:00
  • 9b41f31a66 Use 32x32 black image for VLM server warmup and bring glm4.1v back to UT (#13222) Binyao Jiang 2025-11-13 14:21:43 -08:00
  • 9db3add319 Update GDN causal conv1d cuda kernel - prepare for new changes (#13188) Binyao Jiang 2025-11-13 14:09:47 -08:00
  • c9e5799bd1 [router][grpc] Refine docs in minimax_m2 to match other parsers (#13218) Chang Su 2025-11-13 13:47:24 -08:00
  • 2966367a31 [sgl-kernel] support custom fp8 flashmla kernel (#13087) Fan Yin 2025-11-14 04:45:21 +08:00
  • 8779100768 Fix wrong running_bs in priority scheduling (#13142) dtc 2025-11-14 03:56:38 +08:00
  • dd192a55f4 [Feature] Enable CUDA graph for PD-Multiplexing. (#11595) ykcombat 2025-11-14 03:39:40 +08:00
  • bfe638f7e8 Fix broken Markdown formatting in DeepEP documentation (#13210) Taishi Nakamura 2025-11-14 04:34:57 +09:00
  • 4ac65e3c7d [PD / HiCache]fix deocde kvcache offload manager memory leak (#12774) huangtingwei 2025-11-14 03:31:53 +08:00
  • 0779c3d148 docs: update fused MoE config path (#13211) Junlin Zhou 2025-11-14 03:14:01 +08:00
  • 5c2d72ba63 Bump actions/download-artifact from v4 to v6 for B200 workers (#13220) Kangyan-Zhou 2025-11-13 10:23:58 -08:00
  • 9bd511a582 Add model validation for all GPU runners to prevent cache corruption (#13171) alisonshao 2025-11-13 10:16:41 -08:00
  • 85b8c5c4cd Set max parallel for 1-gpu runner (#13215) Liangsheng Yin 2025-11-14 01:36:43 +08:00
  • c8b7516fe8 Fix accept rate in speculative decoding metrics (#13212) neo 2025-11-14 00:54:50 +08:00
  • 67e9d287ee [Quantization] Support Quark Dense + MoE FP8 & FP8 PTPC (#10485) Bowen Bao 2025-11-13 08:16:00 -08:00
  • e7e89349c9 Enable Flashinfer TRTLLM-GEN-MoE FP8 blockwise kernel for Qwen3-Next on Blackwell (#12543) Sam 2025-11-13 19:44:44 +08:00
  • aead0ef5e5 [FEAT][ROCM] enable fused shared expert for Rocm (#12201) Ling Zhang 2025-11-13 16:41:40 +08:00
  • 6664083522 Replace [silu_and_mul_]scaled_fp4_group_quant by Flashinfer equivalent (#12376) Shu Wang 2025-11-13 02:26:00 -06:00
  • c2d69e8b56 fix: display served_model_name in /v1/models (#13155) Edmund Suen 2025-11-13 16:07:55 +08:00
  • 4c1e909a80 [router] minmax-m2 xml tool parser (#13148) Simo Lin 2025-11-13 16:24:54 +09:00
  • 2bb0317e19 Remove enable_dp_attention in deepseek nightly tests (#13190) Kangyan-Zhou 2025-11-12 22:58:57 -08:00
  • 86255f27b4 Revert "fallback to triton mm_persistent kernel when deepGemm fail" (#13178) fzyzcjy 2025-11-13 14:03:35 +08:00
  • e4b2937017 [AMD] Add AITER Custom All-Reduce (#13102) Hubert Lu 2025-11-12 21:53:44 -08:00
  • 7a8524b444 [Ascend] add npu synchronize (#13154) Mayyyy 2025-11-13 11:03:32 +08:00
  • 909d0d382f docker: fix build error in Dockerfile.diffusion (#12975) Edmund Suen 2025-11-13 09:58:16 +08:00
  • e42df37df0 Remove EBNF Composer (#13163) Tejesh Anand 2025-11-12 17:55:30 -08:00
  • 6d21392b0e [router] remove worker url requirement (#13172) Simo Lin 2025-11-13 10:32:58 +09:00
  • a1cb717d0b Opt kimi_k2_thinking biased topk module (#13150) Xiaoyu Zhang 2025-11-13 08:52:06 +08:00
  • c94564912e Add RequestMetricsExporter utility to export request-level metrics (#10973) Scott Lee 2025-11-12 16:33:37 -08:00
  • c4b74c1db2 [Ascend] LoRA: adding Ascend LoRA backend with using kernels from sgl_kernel_npu (#12288) Vladimir Serov 2025-11-13 03:20:34 +03:00
  • 7aa443903d Fix nan in global scaling factor for large scale nvfp4 EP (#13162) Shu Wang 2025-11-12 17:32:23 -06:00
  • 4eda9969e8 [DeepseekV32]: use _concat_mla_absorb_q_general to replace torch.cat (#12215) bingps 2025-11-13 07:20:01 +08:00
  • 03a7e6f4db Add job and runner failure monitor workflow for CI (#13104) Douglas Yang 2025-11-12 14:04:50 -08:00
  • 2cdde3d46d [router] Fix Flaky test_circuit_breaker_opens_and_recovers (#13164) Xinyue Zhang 2025-11-12 12:07:42 -08:00
  • 401ed0c594 [router] Add comprehensive validation to Responses API (#13127) Keyang Ru 2025-11-12 11:19:27 -08:00
  • 4ef4390540 bugfix: multi-model routing for /generate api (#12979) Siyuan Chen 2025-11-13 03:13:27 +08:00
  • d646cf6347 [Auto Sync] Update test_deterministic.py (20251112) (#13128) Lianmin Zheng 2025-11-12 11:12:23 -08:00
  • 4edb240112 Fuse routed_scaling_factor to fused_marlin_moe (#12998) Ke Bao 2025-11-13 00:18:58 +08:00
  • 706502ff6c [VLM] Support PP for Qwen2.5-VL (#13075) Yuan Luo 2025-11-12 23:18:44 +08:00
  • c2e56dadb2 [Ascend] torch_npu.npu_mrope for MRotaryEmbedding (#10907) Makcum888e 2025-11-12 16:45:54 +03:00
  • 9c1c5c6d7d [ngram] use SGLANG_NGRAM_FORCE_GREEDY_VERIFY to control verify method (#13153) Zhihao Zhang 2025-11-12 21:07:33 +08:00
  • 2d531946db [Ascend][feature] support L1+ L2 radixcache on ascend (#12214) khalilzhk 2025-11-12 20:45:24 +08:00
  • ebaf86d441 chore: bump SGLang version to 0.5.5.post2 (#13129) sglang-bot 2025-11-12 20:35:20 +08:00
  • 8359f185a1 [Feature] Propagate Trace Headers into Root Span for OpenTelemetry Cross-Service Context (#10808) zhanghaotong 2025-11-12 20:28:59 +08:00
  • 5324f37ab3 [Ascend]adapt enable-profile-cuda-graph for NPU (#12617) ronnie_zheng 2025-11-12 14:57:24 +03:00
  • 2864c49fbc [RPC] Fix handle_rpc_request with **recv_req.parameters (#7906) Charlie Ruan 2025-11-12 03:29:12 -08:00
  • d26ec39ff2 Update aiter to v0.1.7.post1 (#13149) sogalin 2025-11-12 19:09:44 +08:00
  • 9c546bfd6d Fix the Wrong Return Type of Scheduler.recv_requests (#7886) Yikai Zhang 2025-11-12 04:54:56 -06:00
  • c9b581644d Dump total_throughput to output-file in bench_serving.py (#9790) Rohan Potdar 2025-11-12 04:48:28 -06:00