Commit Graph

  • f4a8987f69 Update amd docker and nightly models. (#6687) Sai Enduri 2025-05-28 00:08:08 -07:00
  • 41ba767f0c feat: Add warnings for invalid tool_choice and UTs (#6582) Chang Su 2025-05-27 16:53:19 -07:00
  • f127355a30 Add batch test for draft extend (#6672) Ke Bao 2025-05-28 07:32:05 +08:00
  • bdb962d755 fix(tool call): Fix tool_index in PythonicDetector and issues with mixed output in non-streaming (#6678) Chang Su 2025-05-27 16:18:42 -07:00
  • 0b9557fcd7 Disable compiling arch below sm_90 in aarch64 by default (#6380) Qiaolin Yu 2025-05-27 18:50:02 -04:00
  • 87068b5cc7 Support gathering expert distribution details (#6665) fzyzcjy 2025-05-28 06:32:59 +08:00
  • a564e001b5 Fix DeepEP error in Qwen 3 MoE models (#6673) fzyzcjy 2025-05-28 06:12:54 +08:00
  • 2103b80607 [CI] update verlengine ci to 4-gpu test (#6007) Junrong Lin 2025-05-28 05:32:23 +08:00
  • e806f708c9 [PD] Make bootstrap code common between NIXL and Mooncake (#6473) Trevor Morris 2025-05-27 12:47:38 -07:00
  • fa6723f08f Revert "fix communicator for non-dp lm head (#6662)" (#6677) Yineng Zhang 2025-05-27 12:22:59 -07:00
  • 673ff668f7 Speed up expert location update (#6661) fzyzcjy 2025-05-28 01:00:09 +08:00
  • 447be24228 Fix OOM when updating expert locations (#6660) fzyzcjy 2025-05-28 00:59:53 +08:00
  • 183d9f969c DeepSeek: enable none block-quant FP8 quantizations (#6638) HAI 2025-05-27 09:06:40 -07:00
  • 631950280a Support EAGLE draft extend CUDA graph (#6606) Ke Bao 2025-05-27 17:35:17 +08:00
  • a3d7f4b673 fix communicator for non-dp lm head (#6662) Cheng Wan 2025-05-27 02:31:12 -07:00
  • b18416fbf8 Fix qwen3 tbo/dp-lm-head (#6652) Yi Zhang 2025-05-27 15:38:27 +08:00
  • ce9d690ef4 fix: fix nightly test from updating transformers (#6658) Mick 2025-05-27 15:28:11 +08:00
  • bdaefbbfbd Add environment flag for disabling message queue broadcaster (#6403) Baizhou Zhang 2025-05-26 22:32:41 -07:00
  • 45a31a82e4 docs: Update documentation to reflect xgrammar as default grammar backend (#6601) Vincent Zhong 2025-05-27 01:29:13 -04:00
  • 1aa0fbf416 Add note to add supported model to documentation (#6640) Brayden Zhong 2025-05-27 01:18:46 -04:00
  • 7a0bbe6a64 update toc for doc and dockerfile code style format (#6450) linzhuo 2025-05-27 13:05:11 +08:00
  • ae33584235 [Bugfix]: Fix call for function_call_parser.multi_format_detector in adapter.py (#6650) Chang Su 2025-05-26 21:57:10 -07:00
  • 477a101cbd Refactor LoRA handling to support adapter tensors in fused format (#6585) Lifu Huang 2025-05-26 21:51:54 -07:00
  • 1a8f5f6836 Super tiny rename environment variable (#6648) fzyzcjy 2025-05-27 12:01:16 +08:00
  • 32cd707002 Support TP in attention for two batch overlap (#6634) fzyzcjy 2025-05-27 11:28:12 +08:00
  • ebd1ed49d4 Tiny refactor communicator (#6646) fzyzcjy 2025-05-27 11:24:17 +08:00
  • f77da69964 chore: upgrade mooncake-transfer-engine (#6643) Yineng Zhang 2025-05-26 20:01:30 -07:00
  • d6864ce6d6 [New Model] Devstral support (#6547) Xinyuan Tong 2025-05-26 19:27:48 -07:00
  • 755a36614b fix: added "\n" to qwen25 tool parser structural tags (#6631) Shi Shuai 2025-05-27 10:25:45 +08:00
  • 79a39ac0cc follow-up: move Idefics2 to a shared location to eliminate unexpected dependency. (#6603) Lifu Huang 2025-05-26 19:23:59 -07:00
  • 3ce94f71f9 [PD] Handle P/D failure and reconnect without affecting other instances (#6263) shangmingc 2025-05-27 10:21:01 +08:00
  • ca95556c76 Tiny fix sampler error when prob is not contiguous (#6639) fzyzcjy 2025-05-27 10:19:08 +08:00
  • eb8f02dd87 Update nightly thresholds and dependencies. (#6635) Sai Enduri 2025-05-26 11:44:13 -07:00
  • 0ca3e56802 Tiny fix missing expert location dispatch info (#6620) fzyzcjy 2025-05-26 23:58:31 +08:00
  • 5c7aa00976 Fix EPLB algorithm fail to run when using 3 nodes for prefill (#6629) fzyzcjy 2025-05-26 23:43:24 +08:00
  • fe386acae6 Automatically configure for EPLB-related args (#6628) fzyzcjy 2025-05-26 23:42:49 +08:00
  • 14d1075f2c fix qwen3moe eplb prefill bug (#6617) Yi Zhang 2025-05-26 17:15:21 +08:00
  • 006ead9dcb [FA][Test] Fix Sparse FA test (#6306) Brayden Zhong 2025-05-26 04:27:48 -04:00
  • 0d503090aa Supported precomputed feature for Kimi VL (#6599) Lifu Huang 2025-05-26 01:24:13 -07:00
  • 501efc3d36 Tiny fix CI (#6611) fzyzcjy 2025-05-26 14:36:34 +08:00
  • f9bab3d591 qwen3moe support two batch overlap (#6598) Yi Zhang 2025-05-26 14:08:16 +08:00
  • 16f69b1f65 feat: Improve Mistral and Qwen25 function call parsing (#6597) Chang Su 2025-05-25 23:07:23 -07:00
  • 65f091310c refactor qwen moe code, use communicator to support tp+dp (#6581) Yi Zhang 2025-05-26 14:01:10 +08:00
  • fc419b62e8 Revert "Tiny fix lint CI does not trigger on master (#6609)" (#6610) Yineng Zhang 2025-05-25 22:52:34 -07:00
  • 7eb9d8e594 chore: upgrade transformers 4.52.3 (#6575) Yineng Zhang 2025-05-25 22:49:58 -07:00
  • 84147254c9 Tiny fix lint CI does not trigger on master (#6609) fzyzcjy 2025-05-26 13:47:03 +08:00
  • 6bebef60a7 Support accurate length control for bench serving (#6594) fzyzcjy 2025-05-26 13:46:23 +08:00
  • 25be63d0b2 Auto handle PD disaggregation in bench_serving (#6587) fzyzcjy 2025-05-26 13:41:27 +08:00
  • d502dae0f0 Tiny change killall_sglang.sh (#6596) fzyzcjy 2025-05-26 13:36:51 +08:00
  • 93e53f6e0b Logging and minor fixes to two batch overlap and EPLB (#6595) fzyzcjy 2025-05-26 13:36:40 +08:00
  • a191a0e47c Improve performance of two batch overlap in some imbalanced cases (#6593) fzyzcjy 2025-05-26 13:36:18 +08:00
  • 8c7279c24e Fix profiling will crash the server when using num_steps (#6586) fzyzcjy 2025-05-26 13:36:02 +08:00
  • 0ca1811715 Support fake perfectly balanced EP dispatch algorithm (#6571) fzyzcjy 2025-05-26 13:35:51 +08:00
  • 2c3a6fe1de Fix bench_serving does not support changing warmup requests (#6439) fzyzcjy 2025-05-26 13:35:36 +08:00
  • 8b33d8df90 [PD] Fix prefill_servers in mini_lb (#6527) wangxiyu191 2025-05-26 10:38:41 +08:00
  • e235be16fe Fix some issues with current docs. (#6588) simveit 2025-05-25 19:04:34 +02:00
  • 5ccf8fe1a0 Hint users when weight update timeouts (#6570) fzyzcjy 2025-05-26 00:13:17 +08:00
  • 3f23d8cdf1 added support for tied weights in qwen pipeline parallelism (#6546) Shenggui Li 2025-05-25 15:00:56 +08:00
  • 1a39979993 Sgl-router Prometheus metrics endpoint and usage track metrics (#6537) Chao Yang 2025-05-24 22:28:15 -07:00
  • 022012aae8 Support Phi-4 Multi-Modal (text + vision only) (#6494) Lifu Huang 2025-05-24 21:43:38 -07:00
  • 681e7af32b [OAI] Support non-normalized logprobs in OpenAI server (#5961) Chang Su 2025-05-24 21:35:55 -07:00
  • 681fdc264b Refactor vlm embedding routine to use precomputed feature (#6543) Xinyuan Tong 2025-05-24 18:39:21 -07:00
  • 0d47788025 Support overlapping two batches (#4068) fzyzcjy 2025-05-25 08:39:07 +08:00
  • f456037396 Utilize static dispatching for communicator (#6577) fzyzcjy 2025-05-25 08:34:35 +08:00
  • b2388433be Add back DeepSeek non-TBO branches (#6578) fzyzcjy 2025-05-25 08:34:00 +08:00
  • a38376fa99 Refactor attention into multiple stages (#6477) fzyzcjy 2025-05-25 08:33:25 +08:00
  • 7a5e6ce1cb Fix GPU OOM (#6564) kk 2025-05-25 07:38:39 +08:00
  • 24c035f2e3 Temporarily disable MI325x 8 gpu testing. (#6576) Sai Enduri 2025-05-24 16:37:22 -07:00
  • 7e257cd666 chore: bump v0.4.6.post5 (#6566) Yineng Zhang 2025-05-24 00:48:05 -07:00
  • c4831e2fcf Fix accuracy is zero when enabling moe-dense-tp-size as in large scale EP (#6567) fzyzcjy 2025-05-24 15:27:10 +08:00
  • 2e37fa07ba [FIX]remove ServerArgs duplicate code (#6485) Neo 2025-05-24 13:54:41 +08:00
  • 2d831c6ef9 [PD] Support structured output (#6560) Byron Hsu 2025-05-23 21:49:00 -07:00
  • ed0c3035cd feat(Tool Calling): Support required and specific function mode (#6550) Chang Su 2025-05-23 21:00:37 -07:00
  • e6f113569e support eplb for qwen3 (#6533) Yi Zhang 2025-05-24 09:31:30 +08:00
  • 7b02c32679 [Bugfix](gemma3_mm): handle flatten_batch constraint for multiple images (#6562) Chang Su 2025-05-23 18:11:54 -07:00
  • fefa19fec0 Update cmdline --enable-dp-attention help string for Qwen 2/3 Moe models. (#6524) miter 2025-05-24 06:20:21 +08:00
  • 9c574585b3 fix: remove content=none test when tool called (#6347) Shi Shuai 2025-05-24 06:12:55 +08:00
  • 8233cc10fd [PD] Support logprob & Add failure test (#6558) Byron Hsu 2025-05-23 14:29:20 -07:00
  • 1b2e8f76d9 [2/2] Support Qserve (#6521) HandH1998 2025-05-24 03:39:18 +08:00
  • d2e0881a34 [PD] support spec decode (#6507) Byron Hsu 2025-05-23 12:03:05 -07:00
  • 2f42749184 Fix topk inference performance reduce (#6474) Li Hui 2025-05-23 17:58:31 +08:00
  • d8189660a9 Update sgl-kernel UTs for activation/topk/norm/rope kernels (#6452) YanbingJiang 2025-05-23 17:03:15 +08:00
  • 3ded6235c9 Add fp8 fused_experts kernel for CPU in sgl-kernel and add UT (#6404) Chunyuan WU 2025-05-23 17:01:55 +08:00
  • 4ba1eea83f Add fp8 qkv_proj_with_rope kernel for CPU in sgl-kernel and add UT (#6493) blzheng 2025-05-23 15:14:46 +08:00
  • 4685fbb888 [VLM] Support chunk prefill for VLM (#6355) Chang Su 2025-05-22 20:32:41 -07:00
  • 0a4fc73b48 [PD] Fix failure abort (#6535) Byron Hsu 2025-05-22 20:32:03 -07:00
  • a6970a17f3 misc: fix accept_length (#6536) Yineng Zhang 2025-05-22 14:27:10 -07:00
  • a6ae3af15e Support XiaomiMiMo inference with mtp (#6059) ryang 2025-05-23 05:14:49 +08:00
  • 0b07c4a99f chore: upgrade sgl-kernel v0.1.4 (#6532) Yineng Zhang 2025-05-22 13:28:16 -07:00
  • fc0e3b9174 Support qwen3 deepep (#6120) lukec 2025-05-23 02:04:45 +08:00
  • d71f3f0a2a chore: bump sgl-kernel v0.1.4 (#6522) Yineng Zhang 2025-05-22 09:47:42 -07:00
  • 58f10679e1 Fix missing http status import for PD failure handler (#6520) shangmingc 2025-05-22 15:23:54 +08:00
  • 7a80f56513 Support dynamically rebalancing experts using EPLB (#6469) fzyzcjy 2025-05-22 14:13:21 +08:00
  • 9484eba4ad Support logging expert balancedness metrics (#6482) fzyzcjy 2025-05-22 14:05:33 +08:00
  • e9feb48838 [RL] Remove the w13 weight_scale and input_scale for UnquantizedEPMoE… (#6308) Zilin Zhu 2025-05-22 13:03:15 +08:00
  • fc992a09f9 Support updating expert locations dynamically (#6388) fzyzcjy 2025-05-22 12:59:33 +08:00
  • 121f92c583 Add main for merge state tests (#6492) Yuan Luo 2025-05-22 12:56:25 +08:00
  • 3bde101099 [PD] Abort request if transfer fails (#6504) Byron Hsu 2025-05-21 21:44:25 -07:00
  • 7513558074 [PD] Add doc and simplify sender.send (#6019) Byron Hsu 2025-05-21 21:22:21 -07:00
  • 4d643f6c7a [1/2] Support Qserve (#6457) HandH1998 2025-05-22 10:48:59 +08:00