Commit Graph

  • 5f91c82526 [Feature] Support Flashinfer fmha on Blackwell (#6930) Jianan Ji 2025-06-06 15:57:50 -04:00
  • b819381fec AITER backend extension and workload optimizations (#6838) HAI 2025-06-05 23:00:18 -07:00
  • 562f279a2d [CPU] enable CI for PRs, add Dockerfile and auto build task (#6458) Zaili Wang 2025-06-06 04:43:54 +08:00
  • 8b2474898b bugfix(OAI): Fix image_data processing for jinja chat templates (#6877) Chang Su 2025-06-05 13:37:01 -07:00
  • 0df6765c83 [CUTLASS-FP4-MOE] Introduce CutlassMoEParams class for easy initialization of Cutlass Grouped Gems Metadata (#6887) Pavani Majety 2025-06-05 13:13:14 -07:00
  • 35b65cf0ca Use deepgemm instead of triton for fused_qkv_a_proj_with_mqa (#6890) fzyzcjy 2025-06-06 02:37:05 +08:00
  • dd1012fcbe [PD] Fix potential perf spike caused by tracker gc and optimize doc (#6764) shangmingc 2025-06-06 01:56:02 +08:00
  • 44aab7f91c oai: fix openAI client error with single request via batch api (#6170) Ravi Theja 2025-06-05 15:51:47 +05:30
  • 43baba649e [EP] Add cuda kernel for moe_ep_post_reorder (#6837) Yuan Luo 2025-06-05 15:33:47 +08:00
  • 0166403c20 Support Blackwell DeepEP docker images (#6868) fzyzcjy 2025-06-05 15:07:53 +08:00
  • bcf66ef3e1 Tiny allow profiler API to auto create directory (#6865) fzyzcjy 2025-06-05 15:07:03 +08:00
  • 0de5e7d40f Support layerwise rebalancing experts (#6851) fzyzcjy 2025-06-05 15:05:52 +08:00
  • 72a110f664 Tiny update error hints (#6846) fzyzcjy 2025-06-05 15:05:28 +08:00
  • 5aff1e9392 Fix Qwen3MoE missing token padding optimization (#6820) fzyzcjy 2025-06-05 15:04:59 +08:00
  • 8e3797be1c support 1 shot allreduce in 1-node and 2-node using mscclpp (#6277) zyksir 2025-06-05 13:11:24 +08:00
  • 4474eaf552 Support LoRA in TestOpenAIVisionServer and fix fused kv_proj loading bug. (#6861) Lifu Huang 2025-06-04 22:08:30 -07:00
  • 499f5e620c Fix one missing arg in DeepEP (#6878) Cheng Wan 2025-06-04 19:14:47 -07:00
  • 81964328b7 Set num_fused_shared_experts as num_shared_experts when shared_experts fusion is not disabled (#6736) Cheng Wan 2025-06-04 15:53:22 -07:00
  • f0f84975f4 feat: add dp-rank to KV events (#6852) ishandhanani 2025-06-04 15:29:34 -07:00
  • 3f1e433903 Decoder-only Scoring API (#6460) Chanh Nguyen 2025-06-04 14:14:54 -07:00
  • cf9815ba69 [Refactor] Multimodal data processing for VLM (#6659) Xinyuan Tong 2025-06-04 11:22:33 -07:00
  • bd75690f4e fix ep_moe_reorder kernel bugs (#6858) Xiaoyu Zhang 2025-06-04 19:13:59 +08:00
  • 180ff5eecc [fix] recover auto-dispatch for rmsnorm and rope (#6745) JieXin Liang 2025-06-04 12:44:20 +08:00
  • 37f1547587 [FEAT] Add transformers backend support (#5929) Marc Sun 2025-06-04 06:05:29 +02:00
  • 8a5480528d [Refactor] Rename n_share_experts_fusion as num_fused_shared_experts (#6735) Cheng Wan 2025-06-03 17:48:24 -07:00
  • b6d0ce9f78 Minor add metrics to expert location updater (#6816) fzyzcjy 2025-06-03 14:59:11 +08:00
  • 0ea330ca34 Fix wrong weight reference in dynamic EPLB (#6818) fzyzcjy 2025-06-03 14:26:04 +08:00
  • 27e327b415 fix new_page_count_next_decode (#6671) pansicheng 2025-06-03 13:48:52 +08:00
  • ff00895c46 Add CPU optimized kernels for topk and rope fusions (#6456) jianan-gu 2025-06-03 08:37:34 +08:00
  • ff91474825 [Router] Fix k8s Service Discovery (#6766) Arthur Cheng 2025-06-02 16:57:23 -07:00
  • eb38c7d1ca [1/2] Add Kernel support for Cutlass based Fused FP4 MoE (#6093) Pavani Majety 2025-06-02 13:48:03 -07:00
  • df7f61ee7d Speed up rebalancing when using non-static dispatch algorithms (#6812) fzyzcjy 2025-06-03 02:18:17 +08:00
  • ef21729c1d Fix profiles do not have consistent names (#6811) fzyzcjy 2025-06-03 02:17:22 +08:00
  • f5159315b2 Add simple utility to dump tensors for debugging (#6815) fzyzcjy 2025-06-03 02:15:31 +08:00
  • 6d7b6696d4 Tiny fix EPLB assertion about rebalancing period and recorder window size (#6813) fzyzcjy 2025-06-03 02:13:33 +08:00
  • 6376b632eb Tiny log prefill time (#6780) fzyzcjy 2025-06-03 01:28:27 +08:00
  • e05e29d178 Refactor CustomOp to avoid confusing bugs (#5382) fzyzcjy 2025-06-03 01:27:36 +08:00
  • a2cb5913a0 Add draft extend CUDA graph for flashinfer backend (#6805) Ke Bao 2025-06-02 16:51:26 +08:00
  • 55444ed667 [EP] Add cuda kernel for moe_ep_pre_reorder (#6699) Yuan Luo 2025-06-02 11:49:01 +08:00
  • 20fd53b8f6 Correctly abort the failed grammar requests & Improve the handling of abort (#6803) Lianmin Zheng 2025-06-01 19:00:07 -07:00
  • 6a47b73024 Remove contiguous before Flashinfer groupwise fp8 gemm (#6804) Baizhou Zhang 2025-06-01 18:30:54 -07:00
  • c429919def misc: cache is_hopper_arch (#6799) Wenxuan Tan 2025-06-01 17:28:31 -05:00
  • 1da8d23051 chore: update blackwell docker (#6800) Yineng Zhang 2025-06-01 13:37:40 -07:00
  • 2f7420bc84 [Feat] Enable PDL automatically on Hopper architecture (#5981) Huapeng Zhou 2025-06-01 12:30:17 -07:00
  • c6a0cacc35 Update CI tests for Llama4 models (#6421) Ravi Theja 2025-06-01 09:22:15 +05:30
  • 0a9bfc20ab [Minor] Always append newline after image token when parsing chat message (#6797) Lifu Huang 2025-05-31 20:50:33 -07:00
  • 34c63731fc chore: upgrade sgl-kernel v0.1.5 (#6795) Yineng Zhang 2025-05-31 18:32:00 -07:00
  • 2d72fc47cf Improve profiler and integrate profiler in bench_one_batch_server (#6787) Lianmin Zheng 2025-05-31 15:53:55 -07:00
  • b520d02888 chore: bump sgl-kernel v0.1.5 (#6794) Yineng Zhang 2025-05-31 14:54:00 -07:00
  • 7dc0e39442 Bump torch to 2.7.0 (#6788) Qiaolin Yu 2025-05-31 17:43:12 -04:00
  • fb507b7b10 [FIX] mmmu bench serving result display error (#6525) (#6791) Yikai Zhang 2025-06-01 04:48:06 +08:00
  • f90945c45a fix(PD-disaggregation): Can not get local ip (#6792) storyicon 2025-06-01 04:47:14 +08:00
  • 094fbdacd5 Fix incorrect LoRA weight loading for fused gate_up_proj (#6734) Lifu Huang 2025-05-31 13:41:44 -07:00
  • 888cb175a6 Add intel_amx backend for Radix Attention for CPU (#6408) YanbingJiang 2025-05-31 12:37:42 +08:00
  • e39bca0756 ci: relax test_function_call_required (#6786) Chang Su 2025-05-30 19:18:42 -07:00
  • a2bb856543 Temporarily lower mmlu threshold for triton sliding window backend (#6785) Jianan Ji 2025-05-30 21:40:50 -04:00
  • ced3c07afe Support token-level quantization for EP MoE (#6782) Cheng Wan 2025-05-30 17:26:30 -07:00
  • f18b068f15 feat(tool call): Enhance Llama32Detector for improved JSON parsing in non-stream (#6784) Chang Su 2025-05-30 17:05:17 -07:00
  • 4fac524b14 update llama4 chat template and pythonic parser (#6679) Chao Yang 2025-05-30 17:01:22 -07:00
  • b581b22504 Fix one bug in the grouped-gemm triton kernel (#6772) Cheng Wan 2025-05-30 01:42:08 -07:00
  • 69dd878b51 Fix shared experts fusion error (#6289) Li Hui 2025-05-30 16:16:11 +08:00
  • 22630ca242 Support sliding window in triton backend (#6509) Jianan Ji 2025-05-30 04:11:53 -04:00
  • d279d4990c Fix aiohttp 'Chunk too big' in bench_serving (#6737) Yuhong Guo 2025-05-30 15:50:36 +08:00
  • 6cb00c6398 [PD] Optimize time out logic and add env var doc for mooncake (#6761) shangmingc 2025-05-30 15:45:02 +08:00
  • 62cac2c43a Update DeepSeek-R1-0528 function call chat template (#6765) Xu Wenqing 2025-05-30 15:42:57 +08:00
  • 2c3b71d678 Improve EPLB logical to physical dispatch map (#6727) fzyzcjy 2025-05-30 10:23:54 +08:00
  • 51cdd81f97 [fix][RL] Fix DeepSeekV3ForCausalLM.post_load_weights for multiple update weight (#6265) Zilin Zhu 2025-05-30 07:28:10 +08:00
  • 73def253b5 Fix mem_fraction_static for AMD CI (#6748) Baizhou Zhang 2025-05-29 12:37:30 -07:00
  • d9d35def3d [test] add ut and bm for get_last_loc (#6746) JieXin Liang 2025-05-30 02:47:21 +08:00
  • 6df81e8a39 Support tuning DeepEP configs (#6742) fzyzcjy 2025-05-29 23:12:22 +08:00
  • 3ab7d9b55e Support picking variants of EPLB algorithms (#6728) fzyzcjy 2025-05-29 23:12:01 +08:00
  • 7e5071c92a Super tiny enable sole usage of expert distribution metrics and update doc (#6680) fzyzcjy 2025-05-29 23:11:38 +08:00
  • 78689d3393 PD Rust LB (PO2) (#6437) Liangsheng Yin 2025-05-29 20:50:10 +08:00
  • 1dc6864f17 [PD] Support completion endpoint (#6729) shangmingc 2025-05-29 16:26:18 +08:00
  • 485a023bd8 refactor apply_w8a8_block_fp8_linear in fp (#6545) ChangyiYang 2025-05-29 00:15:11 -07:00
  • 7e41290082 Add draft extend CUDA graph for Triton backend (#6705) Ke Bao 2025-05-29 15:13:07 +08:00
  • c673727e0e refactor(tool call): Fix BaseFormatDetector tool_index issue and refactor parse_streaming_increment (#6715) Chang Su 2025-05-29 00:08:45 -07:00
  • f4d4f93928 Add DeepSeek-R1-0528 function call chat template (#6725) Xu Wenqing 2025-05-29 15:05:07 +08:00
  • f2bd3515fb Tune memory arguments on B200 (#6718) Baizhou Zhang 2025-05-29 00:03:22 -07:00
  • c459536b0f [PD] bug fix: Update status if nixl receiver send a a dummy req. (#6720) dongmao zhang 2025-05-29 00:01:56 -07:00
  • 535c838674 [fix] more mem for draft_extend cuda_graph (#6726) JieXin Liang 2025-05-29 14:25:18 +08:00
  • 2163586e63 [feat] triton kernel for get_last_loc (#6676) JieXin Liang 2025-05-29 14:10:28 +08:00
  • e06b076105 Fix PP for Qwen3 MoE (#6709) iLeGend 2025-05-29 14:06:18 +08:00
  • 844a8f42c7 Fix LoRA bench (#6719) Wenxuan Tan 2025-05-28 18:38:55 -05:00
  • 791b3bfabb [Feature] Support Flashinfer fp8 blockwise GEMM kernel on Blackwell (#6479) Baizhou Zhang 2025-05-28 16:03:43 -07:00
  • 31589e177e Speed up when having padding tokens two-batch overlap (#6668) fzyzcjy 2025-05-29 07:00:58 +08:00
  • ae6a5b2950 Minor refactor two-batch overlap (#6682) fzyzcjy 2025-05-29 06:54:17 +08:00
  • 4839999b76 Overlap two kernels in DeepSeek with communication (#6711) fzyzcjy 2025-05-29 06:53:51 +08:00
  • 541a985f85 Fuse routed_scaling_factor in DeepSeek (#6710) fzyzcjy 2025-05-29 06:53:37 +08:00
  • 5170b010a6 [PD] Remove Unnecessary Exception Handling for FastQueue.get() (#6712) Hongbo Xu 2025-05-29 02:18:24 +08:00
  • d63e76f735 [CI] Fix setup of disaggregation with different tp (#6706) shangmingc 2025-05-29 02:17:27 +08:00
  • e9fd11c0d1 [Bugfix] Fix ChatCompletion endpoint of mini_lb when stream is set (#6703) shangmingc 2025-05-28 21:33:36 +08:00
  • c7588d593e [Bugfix] Fix slice operation when chunk size mismatch (#6697) shangmingc 2025-05-28 21:15:00 +08:00
  • 6b231325b9 [PD Perf] replace Queue to FastQueue (#6649) ybyang 2025-05-28 16:37:51 +08:00
  • b1c8d4e9f3 [PD] Abort unbootstrapped prefill requests through timeout (#6685) shangmingc 2025-05-28 15:40:54 +08:00
  • c25231c679 [CI] Fix flaky pp single node test (#6689) shangmingc 2025-05-28 15:40:26 +08:00
  • fba03b29e3 [Bugfix] Fix missing abort finish reason for PD with ChatCompletion (#6693) shangmingc 2025-05-28 15:39:46 +08:00
  • 461a730280 fix(deepseekv3): Fix DeepSeekV3Detector tool_index assignment and multi-tool call streaming support (#6655) Chang Su 2025-05-28 00:22:53 -07:00
  • 076103535c fix log_info_on_rank0 error when run benchmark (#6260) Xiaoyu Zhang 2025-05-28 15:20:01 +08:00
  • c087ddd686 Refine pre_reorder_triton_kernel slightly to improve performance (#6627) Yuan Luo 2025-05-28 15:15:23 +08:00