Commit Graph

  • cd7c8a8de6 doc: update developer guide regarding mllms (#6138) Mick 2025-05-14 23:13:13 +08:00
  • 3e350a931e [Bug] Fix accidental logger override caused by internVL. (#6282) Lifu Huang 2025-05-13 23:29:25 -07:00
  • fb71725c98 Fix a bug in schedule_policy (#6276) Ying Sheng 2025-05-13 18:04:00 -07:00
  • 912788c095 perf: optimize local_block_table memory allocation (#6273) Chang Su 2025-05-13 17:18:38 -07:00
  • 0f75b907c6 [CPU] Add CMakeLists.txt for sgl-kernel (#6115) blzheng 2025-05-14 06:30:37 +08:00
  • 16267d4fa7 chore: bump v0.4.6.post4 (#6245) Yineng Zhang 2025-05-13 01:57:51 -07:00
  • 0f5cb8cae1 Enable MI325X AMD CI. (#6259) Sai Enduri 2025-05-13 01:49:33 -07:00
  • 17299f088a [misc] deep_gemm fallback to NVRTC when NVCC not found (#6252) JieXin Liang 2025-05-13 16:41:35 +08:00
  • 5380cd7ea3 model(vlm): pixtral (#5084) Kiv Chen 2025-05-13 00:16:10 -07:00
  • b2e95f62b4 Fix two issues related to --moe-dense-tp-size=1 (#5657) Cheng Wan 2025-05-13 02:51:39 -04:00
  • 1ab14c4c5c [VERL Use Case] Add torch_memory_saver into deps (#6247) Stefan He 2025-05-12 19:09:03 -07:00
  • 3c32895cbe [Llama4] Add docs note about enable multimodal (#6235) Brayden Zhong 2025-05-12 22:05:47 -04:00
  • ac2324c177 Skip the flaky test_stateful_custom_logit_processor (#6251) Lianmin Zheng 2025-05-12 18:29:41 -07:00
  • ef8ec07b2c Support tuning moe for llama 4 model (#6042) fzyzcjy 2025-05-13 06:47:01 +08:00
  • f24fc5b86d fix typo (#6248) Yineng Zhang 2025-05-12 15:45:12 -07:00
  • d18c6b3358 Support incremental streaming of logprob/token_ids between scheduler and detokenizer (#6225) Lianmin Zheng 2025-05-12 14:33:38 -07:00
  • f1c896007a [PD] Add support for different TP sizes per DP rank (#5922) shangmingc 2025-05-13 04:55:42 +08:00
  • 983c663de6 Update AMD nightly deps. (#6241) Sai Enduri 2025-05-12 13:39:20 -07:00
  • f94543d22b chore: add hf_xet dep (#6243) Yineng Zhang 2025-05-12 13:08:40 -07:00
  • e8e18dcdcc Revert "fix some typos" (#6244) Lianmin Zheng 2025-05-12 12:53:26 -07:00
  • bad7c26fdc [PP] Fix init_memory_pool desync & add PP for mixtral (#6223) Ying Sheng 2025-05-12 12:38:09 -07:00
  • 12319a6787 [Docs] Add docs for SGLANG_ and SGL_ environment variables (#6206) Brayden Zhong 2025-05-12 13:45:41 -04:00
  • d738ab52f8 fix some typos (#6209) applesaucethebun 2025-05-12 13:42:38 -04:00
  • 3ee40ff919 [CI] Re-enable pd disaggregation test (#6231) shangmingc 2025-05-13 01:09:12 +08:00
  • 0f334945c6 [CI] Fix PD mooncake dependency error (#6212) shangmingc 2025-05-13 01:08:49 +08:00
  • fba8eccd7e Log if cuda graph is used & extend cuda graph capture to cuda-graph-max-bs (#6201) Lianmin Zheng 2025-05-12 00:17:33 -07:00
  • 7d3a3d4510 Update AMD CI docker to v0.4.6.post3-rocm630. (#6213) Sai Enduri 2025-05-12 00:00:46 -07:00
  • 25c83fff6a Performing Vocabulary Parallelism for LM Head across Attention TP Groups (#5558) Cheng Wan 2025-05-12 02:36:29 -04:00
  • 9f2c9568f0 [doc] add a note for --n-share-experts-fusion args (#6154) Xiaoyu Zhang 2025-05-12 14:18:38 +08:00
  • 3f2702ae51 Fix start_profile does not support with_stack and record_shapes (#6043) fzyzcjy 2025-05-12 14:11:32 +08:00
  • 6ea05950b1 Fix release-docs.yml to not use python 3.9 (#6204) Lianmin Zheng 2025-05-11 16:04:55 -07:00
  • e7dd906c5c Update README.md (#6202) Lianmin Zheng 2025-05-11 14:34:12 -07:00
  • 6e2da51561 Replace time.time() to time.perf_counter() for benchmarking. (#6178) Lifu Huang 2025-05-11 14:32:49 -07:00
  • e9a47f4cb5 Add dev-deepep docker image (#6198) fzyzcjy 2025-05-12 04:17:55 +08:00
  • 03227c5fa6 [CI] Reorganize the 8 gpu tests (#6192) Lianmin Zheng 2025-05-11 10:55:06 -07:00
  • 01bdbf7f80 Improve structured outputs: fix race condition, server crash, metrics and style (#6188) Lianmin Zheng 2025-05-11 08:36:16 -07:00
  • 94d42b6794 [Docs] minor Qwen3 and reasoning parser docs fix (#6032) Adarsh Shirawalmath 2025-05-11 20:52:46 +05:30
  • 69276f619a doc: fix the erroneous documents and example codes about Alibaba-NLP/gme-Qwen2-VL-2B-Instruct (#6199) mlmz 2025-05-11 23:22:11 +08:00
  • 41a645f556 Handle empty input string for embedding models (#5621) Ravi Theja 2025-05-11 20:47:15 +05:30
  • 230106304d chore: upgrade sgl-kernel v0.1.2.post1 (#6196) Yineng Zhang 2025-05-11 07:41:37 -07:00
  • 45b4dcf037 chore: bump sgl-kernel v0.1.2.post1 (#6195) Yineng Zhang 2025-05-11 02:24:10 -07:00
  • 213e8c7dd5 chore: upgrade deepgemm (#6073) Yineng Zhang 2025-05-11 02:17:24 -07:00
  • 41273fd71f fix: handle None multimodal_inputs during merging and filtering batches in disaggregation decode mode (#6169) Yusong Gao 2025-05-11 15:28:21 +08:00
  • e9bebafb19 [fix] remove mixtral from is_fa3_default_architecture (#6191) JieXin Liang 2025-05-11 15:15:54 +08:00
  • 4d1c9db66c feat: support loogle eval (#6190) Yineng Zhang 2025-05-10 23:52:44 -07:00
  • 17c36c5511 [CI] Disabled deepep tests temporarily because it takes too much time. (#6186) Lianmin Zheng 2025-05-10 23:40:50 -07:00
  • 31d1f6e7f4 [PD] Add simple unit test for disaggregation feature (#5654) shangmingc 2025-05-11 13:35:27 +08:00
  • a823c6e834 Remove duplicate IO Struct test (#6180) Emmanuel Ferdman 2025-05-11 07:57:26 +03:00
  • 2ce8793519 Add typo checker in pre-commit (#6179) applesaucethebun 2025-05-11 00:55:00 -04:00
  • de167cf5fa Fix request abortion (#6184) Lianmin Zheng 2025-05-10 21:54:46 -07:00
  • 4319978c73 Fix data parallel perf regression (#6183) Lianmin Zheng 2025-05-10 19:18:35 -07:00
  • 03dd785cd0 Added async_encode method to Engine (#4701) Steven Shimizu 2025-05-10 18:58:40 -07:00
  • 66fc63d6b1 Revert "feat: add thinking_budget (#6089)" (#6181) Yineng Zhang 2025-05-10 16:07:45 -07:00
  • 921e4a8185 [Docs]Delete duplicate content (#6146) Ximingwang-09 2025-05-11 06:02:15 +08:00
  • 9d8ec2e67e Fix and Clean up chat-template requirement for VLM (#6114) XinyuanTong 2025-05-10 09:14:09 -07:00
  • c178abdabc [fix] fix determine_n_share_experts_fusion (#6118) JieXin Liang 2025-05-10 16:19:09 +08:00
  • b29a026e14 KV‑Cache (MHA, MLA): add missing start_layer / end_layer fields to MHATokenToKVPoolHost and MLATokenToKVPoolHost (#6016) Simon (Jiyou) Li 2025-05-10 06:50:06 +08:00
  • 678d8cc987 chore: bump v0.4.6.post3 (#6165) Yineng Zhang 2025-05-09 15:38:47 -07:00
  • d2cb3024f2 fix bug that gpu0 occupies more memory when hicache is turned on (#5778) huangtingwei 2025-05-10 06:36:08 +08:00
  • 1940cdec61 [Bugfix] Fix Llama4 gibberish output with long context and CUDA graph (#6162) Chang Su 2025-05-09 15:33:02 -07:00
  • 63484f9fd6 feat: add thinking_budget (#6089) thyecust 2025-05-09 23:22:09 +08:00
  • dff0ab92eb Update amd nightly concurrency. (#6141) Sai Enduri 2025-05-09 00:02:14 -07:00
  • e30c273bc9 opt flashinfer mla cat (#5822) xu-yfei 2025-05-09 14:17:14 +08:00
  • 0ab3f437ab Cutlass MLA: Disable split kv due to https://github.com/NVIDIA/cutlass/issues/2274 (#6101) Trevor Morris 2025-05-08 18:44:30 -07:00
  • cec98f1034 [Fix] Incorrect Memory Allocation on CUDA:0 by Non-Zero CUDA Processes in TP/DP (#5745) yhyang201 2025-05-09 08:52:26 +08:00
  • 8dc4efd0ab docs: update README (#6132) Yineng Zhang 2025-05-08 15:18:48 -07:00
  • 6578cf27de chore: bump sgl-kernel 0.1.2 (#6131) Yineng Zhang 2025-05-08 15:16:28 -07:00
  • 087751a8f2 Remove unecessary is_fa3_supported check (#6112) Stefan He 2025-05-08 14:45:33 -07:00
  • 911f3ba6f4 upgrade xgrammar to 0.1.19 (#6129) Yixin Dong 2025-05-08 17:42:02 -04:00
  • f6f96b0521 [sgl-kernel] fix: fix cu118 compile error (#6123) PGFLMG 2025-05-09 05:26:51 +08:00
  • 2a936a841e [AMD] switch to custom allreduce regardless of MSCCL setting on ROCm (#6097) Hubert Lu 2025-05-08 13:46:58 -07:00
  • 5e02330137 [perf] dsv3 bmm fallback to bf16 (#5662) JieXin Liang 2025-05-09 02:43:39 +08:00
  • fa7d7fd9e5 [Feature] Add FlashAttention3 as a backend for VisionAttention (#5764) Zhu Chen 2025-05-09 01:01:19 +08:00
  • f1ff736d68 [fix] fix pyproject.toml dependencies (#6119) JieXin Liang 2025-05-08 17:14:36 +08:00
  • acc816d8a2 DeepEP normal support deepgemm-contiguous (#5626) lukec 2025-05-08 16:20:32 +08:00
  • a05bd83a94 Change AMD test threshold (#6091) fzyzcjy 2025-05-08 16:05:52 +08:00
  • cef91b1ed7 [PD] Add control to slow down a server (#5572) fzyzcjy 2025-05-08 16:03:08 +08:00
  • 6450c1228c Tiny refactor weight loading logic (#5232) fzyzcjy 2025-05-08 16:02:56 +08:00
  • b6cf3532b5 Tiny refactor ModelConfig.from_server_args (#5219) fzyzcjy 2025-05-08 16:02:43 +08:00
  • 3b2680a44d Overlap shared expert and routed expert computations (#5121) fzyzcjy 2025-05-08 16:02:32 +08:00
  • 79961afa82 optimize pad operations in fa3 to accelarate 100+us (#6077) Minglei Zhu 2025-05-07 23:40:08 -07:00
  • cfca4e0ed2 adding Triton configs for DeepSeekV3 FusedMoE kernel on Blackwell (#6111) Baizhou Zhang 2025-05-07 23:39:10 -07:00
  • e88dd482ed [CI]Add performance CI for VLM (#6038) XinyuanTong 2025-05-07 19:20:03 -07:00
  • 73600673bb Clean logs for DeepSeek-V3 launching (#6079) Baizhou Zhang 2025-05-07 18:54:50 -07:00
  • 8f508cc77f Update doc for MLA attention backends (#6034) Baizhou Zhang 2025-05-07 18:51:05 -07:00
  • 9bddf1c82d Deferring 8 GPU test (#6102) Cheng Wan 2025-05-07 21:49:58 -04:00
  • 24c13ca950 Clean up fa3 test from 8 gpus (#6105) Stefan He 2025-05-07 18:38:40 -07:00
  • b70957fcf8 [refactor] slightly tidy fp8 module (#5993) JieXin Liang 2025-05-08 08:28:24 +08:00
  • e444c13fb4 feat(engine): add bootstrap parameters to generate methods (dynamo) (#6075) ishandhanani 2025-05-07 10:33:58 -07:00
  • fee37d9e8d [Doc]Fix description for dp_size argument (#6063) Baizhou Zhang 2025-05-07 09:04:22 -07:00
  • c68de47915 Super tiny fix doc (#5233) fzyzcjy 2025-05-07 22:41:50 +08:00
  • 4c7b42424c Hint users DeepEP normal mode is incompatible with CUDA Graph (#5014) fzyzcjy 2025-05-07 22:40:59 +08:00
  • 38053c3372 Fix the timeout for 8 gpu tests (#6084) Lianmin Zheng 2025-05-07 03:13:12 -07:00
  • 00c2c1f08b [Feature] Support for Ascend NPU backend (#3853) Song Zhang 2025-05-07 11:32:53 +08:00
  • cb69194562 feat: add release workflow for SGLang kernels on aarch64 (#6010) Johnny 2025-05-07 04:42:07 +02:00
  • d25398cbc8 fix custom_allreduce namespace (#6039) Xiaoyu Zhang 2025-05-07 10:13:06 +08:00
  • 8a828666a3 Add DeepEP to CI PR Test (#5655) Jinyan Chen 2025-05-07 08:36:03 +08:00
  • aff584fa54 Fix sgl-kernel build on aarch64 platforms (#6062) Qiaolin Yu 2025-05-06 19:10:57 -04:00
  • 6f56614734 chore: upgrade cutlass 3.9.2 (#6004) Yineng Zhang 2025-05-06 13:34:08 -07:00
  • bdd17998e6 [Fix] Fix and rename flashmla CI test (#6045) Baizhou Zhang 2025-05-06 13:25:15 -07:00