Commit Graph

  • c9abd7be01 fix: deepep dockerfile, use pip install deepep. (#5885) Hank Han 2025-05-07 01:55:26 +08:00
  • a3e4e9bf9e Better PD initialization (#5751) Liangsheng Yin 2025-05-07 01:12:57 +08:00
  • 6d4d3bc81d Fix not "import os" (#6057) Liangsheng Yin 2025-05-06 22:06:41 +08:00
  • 5f300141b7 docs: add new blog (#6048) Yineng Zhang 2025-05-06 00:48:10 -07:00
  • 1c05425bcb docs: add Google Cloud Vertex AI in Adoption and Sponsorship (#6047) Yineng Zhang 2025-05-06 00:31:48 -07:00
  • b26cb1c55a Fix problem of large page size with chunked prefill (#6046) Zhiqiang Xie 2025-05-06 00:19:47 -07:00
  • f8e460930a Fix prefill OOM error in the case of large page size (#5081) Zhiqiang Xie 2025-05-05 16:02:55 -07:00
  • 683707c314 [Security][Bug] Prevent binding to all TCP interfaces (#5752) Adarsh Shirawalmath 2025-05-06 00:51:45 +05:30
  • a68ed76682 feat: append more comprehensive fields in messages instead of merely role and content (#5996) mlmz 2025-05-06 02:43:34 +08:00
  • 82653f6622 feat: Add a unified merge_state API (#5428) DefTruth 2025-05-06 01:32:33 +08:00
  • 22da3d978f Fix "Avoid computing lse in Ragged Prefill when there's no prefix match" (#5555) Wenxuan Tan 2025-05-05 12:32:17 -05:00
  • b8559764f6 [Test] Add flashmla attention backend test (#5587) Huapeng Zhou 2025-05-06 01:32:02 +08:00
  • 56f6589ecb [PD] Optimize disaggregation ib device help info (#5781) shangmingc 2025-05-05 13:47:37 +08:00
  • 1232f7e8b7 Update dev container config to support live code sync and improve docker setup guide (#6018) Lifu Huang 2025-05-04 22:33:46 -07:00
  • 3008db9c1a [PD] Allow customizing reserved tokens to avoid KV cache waste (#6002) fzyzcjy 2025-05-05 11:23:15 +08:00
  • 357fb2dba5 fix: fix broadcast_pyobj breaking VerlEngine (#5997) Junrong Lin 2025-05-05 04:15:53 +08:00
  • 95c231e50d Tool Call: Add chat_template_kwargs documentation (#5679) vzed 2025-05-04 16:12:40 -04:00
  • 3042f1da61 Fix flaky issues of lora and add multi batch tests (#5957) Qiaolin Yu 2025-05-04 16:11:40 -04:00
  • 2b63798c7d [Minor] Fix duplicate method definitions in conversation.py (#6012) Lifu Huang 2025-05-04 13:02:53 -07:00
  • bf203cb7a2 [Fix] Suppress dynamo logging when using flashinfer backend with torch compile (#5992) Baizhou Zhang 2025-05-04 09:49:13 -07:00
  • 8ebde73f7d [perf] H100 DeepSeek-V3 fused moe tuned config (#5998) JieXin Liang 2025-05-04 05:02:26 +08:00
  • 6b0fae797a Fix Phi3 serving which was broke by earlier change (#5991) Stefan He 2025-05-03 00:28:47 -07:00
  • 141a459644 fix: only upgrade nccl for cu128 (#5986) Yineng Zhang 2025-05-02 11:31:29 -07:00
  • d8ab60117f Overlap qk norm with two streams (#5977) Ke Bao 2025-05-03 00:26:30 +08:00
  • 6579cd7daf Fix set kv cache multi-stream (#5975) Ke Bao 2025-05-03 00:26:00 +08:00
  • 97ac42b634 [PD] NIXL backend Prefill TP & Decode TP+DP (#5681) Yongtong Wu 2025-05-02 22:14:03 +08:00
  • 1acca3a2c6 FA3 speed up: skip len operation and get batch size directly from forward batch (#5969) Lifu Huang 2025-05-02 00:26:12 -07:00
  • 6ea1e6ac6e Support MMMU benchmark for InternVL (#5968) XinyuanTong 2025-05-02 00:17:21 -07:00
  • 3409aaab32 Support InternVL3 (#5350) xm:D 2025-05-02 13:38:59 +08:00
  • 73dcf2b326 Remove token in token out in Native API (#5967) Chayenne 2025-05-01 21:59:43 -07:00
  • 170d1f218a feat: Refactor DeepSeekV3 function call (#5908) Chang Su 2025-05-01 21:28:57 -07:00
  • 73bc1d00fc Add 1 gpu perf and 2 gpu accuracy tests for AMD MI300x CI. (#5960) Sai Enduri 2025-05-01 20:56:59 -07:00
  • c5645e928f feat: add concurrency evaluation logic in mmmu benchmark (#5782) XinyuanTong 2025-05-01 18:20:08 -07:00
  • d33955d28a Properly return error response in vertex_generate HTTP endpoint (#5956) KCFindstr 2025-05-01 11:48:58 -07:00
  • 6fc175968c Optimize a pad operation to accelerate 25us (#5945) Stefan He 2025-05-01 10:48:55 -07:00
  • ad506a4e6b docs: Fix Qwen model typo (#5944) 江家瑋 2025-05-02 01:23:00 +08:00
  • ebaba85655 Update ci test and doc for MTP api change (#5952) Ke Bao 2025-05-02 00:30:27 +08:00
  • de2faef97e Remove extra contiguous (#5953) Ke Bao 2025-05-02 00:28:46 +08:00
  • 67b7d5b1df [PD] Vectorise group_concurrent_contiguous in NumPy (#5834) Yuan Luo 2025-05-01 22:42:37 +08:00
  • 4322c31e24 Support XiaomiMiMo/MiMo model inference (#5921) ryang 2025-05-01 22:41:13 +08:00
  • 9858113c33 chore: bump v0.4.6.post2 (#5939) Yineng Zhang 2025-04-30 22:04:40 -07:00
  • 8441baad6e fix: update model runner (#5934) Yineng Zhang 2025-04-30 19:49:26 -07:00
  • 256c4c2519 fix: correct stream response when enable_thinking is set to false (#5881) mlmz 2025-05-01 10:44:37 +08:00
  • 9f21e75453 add Thor & Spark (#5915) Johnny 2025-05-01 04:43:40 +02:00
  • 7bcd8b1cb2 Fix lora batch processing when input lora_path contains None (#5930) Qiaolin Yu 2025-04-30 22:42:42 -04:00
  • 11383cec3c [PP] Add pipeline parallelism (#5724) Ying Sheng 2025-04-30 18:18:07 -07:00
  • e97e57e699 Remove unused method calculate_num_image_tokens from qwen2_vl.py (#5783) XinyuanTong 2025-04-30 17:46:59 -07:00
  • 9a6ad8916d chore: upgrade sgl-kernel 0.1.1 (#5933) Yineng Zhang 2025-04-30 16:13:30 -07:00
  • d353d08b4e chore: bump sgl-kernel 0.1.1 (#5932) Yineng Zhang 2025-04-30 14:01:49 -07:00
  • 08acdb5c3d [Feat] Scale up fa3 kernel to sm8x arch (#5912) PGFLMG 2025-05-01 04:59:36 +08:00
  • 2afba1b1c1 Add TP2 MOE benchmarks for AMD. (#5909) Sai Enduri 2025-04-30 11:38:20 -07:00
  • e330f2b86c [qwen3] support qwen3 ep moe (#5917) laixin 2025-05-01 00:15:21 +08:00
  • 3ddf5b9d61 [Misc] use parallel build for cmake in sgl-kernel (#5919) PGFLMG 2025-04-30 23:56:46 +08:00
  • 3cff963335 [fix] kimi-vl test in test_vision_openai_server.py (#5910) JieXin Liang 2025-04-30 14:59:10 +08:00
  • d50e36a79d support vlm benchmark profile (#5905) Yi Zhang 2025-04-30 14:48:27 +08:00
  • 8fefdd32c7 [Feature] add support kimi vl model (#5383) liwenju0 2025-04-30 12:31:19 +08:00
  • 403b855a22 Add sm_120 for blackwell (#5903) zhjunqin 2025-04-30 11:45:24 +08:00
  • 1698e94e67 Add A800 fused moe config for qwen3 235b (#5900) lambert0312 2025-04-30 11:18:11 +08:00
  • 58195dd588 [Fix] Unload lora in HF_Runner if needed (#5899) Qiaolin Yu 2025-04-29 23:17:42 -04:00
  • 799789afed Bump Flashinfer to 0.2.5 (#5870) Baizhou Zhang 2025-04-29 19:50:57 -07:00
  • cc4a80caf6 [PD] Fix Assertion failed: /DeepEP/csrc/kernels/internode.cu:483, condition: ibgda_get_state()->num_rc_per_pe >= num_channels #134 (#5830) ybyang 2025-04-30 10:38:54 +08:00
  • 3c8a52311a Fix check_env script (#5901) lambert0312 2025-04-30 09:54:54 +08:00
  • a043f7f2ab chore: use torch 2.6 for sgl-kernel build (#5898) Yineng Zhang 2025-04-29 17:51:18 -07:00
  • e3a5304475 Add AMD MI300x Nightly Testing. (#5861) saienduri 2025-04-29 17:34:32 -07:00
  • 28b26dbf48 [Bugfix]: fix missing queue_time_start for requests from grammar_queue (#5696) Chang Su 2025-04-29 17:31:44 -07:00
  • 2b06484bd1 feat: support pythonic tool call and index in tool call streaming (#5725) Chang Su 2025-04-29 17:30:44 -07:00
  • e4b6133b78 [fix] relax mem_fraction_static for h200 (#5893) JieXin Liang 2025-04-30 08:01:12 +08:00
  • dd408ee481 Auto set draft model path for MTP (#5793) Ke Bao 2025-04-30 07:25:40 +08:00
  • 9419e75d60 [CI] Add test_function_calling.py to run_suite.py (#5896) Chang Su 2025-04-29 15:54:53 -07:00
  • 2c7dbb7cc2 [FEATURE] Enhance platform compatibility for ARM (#5746) Johnny 2025-04-30 00:06:16 +02:00
  • 9a62191ba7 chore: update CODEOWNERS (#5895) Yineng Zhang 2025-04-29 14:12:04 -07:00
  • ae523675e5 [Doc] Tables instead of bulletpoints for sampling doc (#5841) simveit 2025-04-29 22:49:39 +02:00
  • 5c08aa4958 [Docs] Update docs for Qwen3 and Qwen3MoE (#5836) Adarsh Shirawalmath 2025-04-30 02:18:30 +05:30
  • f4c191a712 chore: update Dockerfile (#5894) Yineng Zhang 2025-04-29 12:55:13 -07:00
  • 771669cbe0 [fix]: PyO3 macOS linking and consolidate on tracing for logging Simo Lin 2025-04-29 11:26:38 -07:00
  • 1468769bde [Misc] add service discovery for sgl router Simo Lin 2025-04-29 10:21:19 -07:00
  • 91dda4cd06 Add A800 fused moe config for qwen3 30b (#5880) lambert0312 2025-04-29 17:02:24 +08:00
  • 8e5a6d3441 [Fix] Fix a bug for flashmla to run R1 model (#5875) pengcuo 2025-04-29 16:03:13 +08:00
  • 8465f035d1 Add qwen3 30b fused moe config (#5859) XinyuanTong 2025-04-29 00:24:00 -07:00
  • 8c0cfca87d Feat: support cuda graph for LoRA (#4115) Qiaolin Yu 2025-04-29 02:30:44 -04:00
  • 2c3ea29476 [Feature] support auto chat template (#4949) woodx 2025-04-29 13:34:18 +08:00
  • 5bb0accbcf cutlass 3.9 supported to improve fp8_blockwise_gemm (#5820) Xiaoyu Zhang 2025-04-29 12:52:36 +08:00
  • 8d463fe351 Cutlass MLA decode - fix dtype error (#5868) Trevor Morris 2025-04-28 21:12:58 -07:00
  • 26fc32d168 [CI] tune the test order to warmup the server (#5860) Lianmin Zheng 2025-04-28 19:27:37 -07:00
  • 1cc326032d simplify fused_moe config logging (#5801) Xiaoyu Zhang 2025-04-29 08:04:54 +08:00
  • 05ee219286 Support max_completion_tokens for OpenAIChatCompletions (#5857) Chang Su 2025-04-28 13:50:13 -07:00
  • dcae1fb2cd chore: bump v0.4.6.post1 (#5845) Yineng Zhang 2025-04-28 12:57:08 -07:00
  • a0251a3fd6 add fused moe config for qwen3moe fp8/bf16 (#5849) Yi Zhang 2025-04-29 02:55:52 +08:00
  • 663037a7a0 feat: update is_fa3_default_architecture (#5854) Yineng Zhang 2025-04-28 11:53:22 -07:00
  • f4a9f60cbd [Fix] Missing bootstrap_port field (#5823) XTY 2025-04-29 02:13:04 +08:00
  • ee71ed8a41 [Feat] QWen-1M context support[1/2]: Update block sparse attention backend utils kernel (#5847) PGFLMG 2025-04-29 02:03:17 +08:00
  • d364b9b0f2 ROCm: update AITER (#5816) HAI 2025-04-28 11:01:20 -07:00
  • 849c83a0c0 [CI] test chunked prefill more (#5798) Lianmin Zheng 2025-04-28 10:57:17 -07:00
  • d73ddeb196 feat: Add fused moe triton config for qwen3-30b-fp8 moe on h20 (#5850) JiLi 2025-04-29 01:49:33 +08:00
  • f48b007c1d [Doc] Recover history of server_arguments.md (#5851) Baizhou Zhang 2025-04-28 10:48:21 -07:00
  • 74cb12a878 [config] qwen3moe_tune_h20 fp8 tp4 (#5846) ybyang 2025-04-29 01:21:06 +08:00
  • c6c6264073 [PD] support pd fake transfer for warmup (#5726) ybyang 2025-04-29 00:33:20 +08:00
  • 92ab0a2055 feat: Add fused moe triton config for qwen3bf16 moe on h20 (#5839) yhyang201 2025-04-29 00:30:59 +08:00
  • e132cba2a8 fused moe triton tuning script support qwen3 (#5842) Xiaoyu Zhang 2025-04-29 00:13:04 +08:00
  • 0045f4b2af feat: Add fused moe triton config for qwen3 moe on h100 (#5833) XinyuanTong 2025-04-28 08:37:13 -07:00