Commit Graph

  • c04a8a820b [fix] fix misusing of is_cuda (#7790) JieXin Liang 2025-07-05 19:02:14 +08:00
  • 6c903611ca Fix incorrect spec_num_draft_tokens in draft_extend (#7757) Cheng Wan 2025-07-05 02:18:16 -07:00
  • 77cfea689d chore: upgrade sgl-kernel v0.2.3 (#7786) Yineng Zhang 2025-07-05 01:55:55 -07:00
  • 8fc910db03 DP Attention with Auto DeepEP Dispatch (#7222) Cheng Wan 2025-07-05 01:54:24 -07:00
  • 75354d9ae9 fix: use nvidia-nccl-cu12 2.27.5 (#7787) Yineng Zhang 2025-07-05 01:28:21 -07:00
  • 4fece12be9 chore: bump sgl-kernel v0.2.3 (#7784) Yineng Zhang 2025-07-05 00:05:45 -07:00
  • c797322280 fix: fix apply_shuffle_mul_sum (#7444) Mick 2025-07-05 14:23:30 +08:00
  • ef8a29c429 Embedding parallel by attn_tp (#7623) Gang Chen 2025-07-05 14:21:56 +08:00
  • 8e9fb43d82 Optimize Hopper CUTLASS FP8 Blockwise Grouped GEMM Kernel in Small K Scenario (#7782) Qi Yuhang 2025-07-05 13:25:49 +08:00
  • 8364608930 add model: qwen2-audio (#7596) Leng Yue 2025-07-04 21:13:10 -07:00
  • da3890e82a [1/n]: add cutlass W4A8 moe kernel for hopper architecture (#7772) SijiaYang 2025-07-05 11:50:12 +08:00
  • cb432f1770 saving hidden_states.clone() (#7705) Cheng Wan 2025-07-04 20:07:42 -07:00
  • 1964c325de [feat] Support EAGLE3 for Qwen (#7745) Ximingwang-09 2025-07-05 10:50:28 +08:00
  • af5647748a [Fix] Alloc return type error (#7778) Caproni 2025-07-05 10:00:40 +08:00
  • af46f299f9 [RL] add pause and continue generation for async rl training (#7419) Zilin Zhu 2025-07-05 09:49:49 +08:00
  • 16a6b1d83a [RL] Add --nccl-port to prevent port conflict (#7418) Zilin Zhu 2025-07-05 09:48:57 +08:00
  • 14229ccf8f Move mem_fraction_static adjustment for multimodal models to server_args.py & Fix session control & Other cleanups (#7748) Lianmin Zheng 2025-07-04 16:33:33 -07:00
  • 975a5ec69c [fix] update bench_speculative.py for compatibility (#7764) Kay Yan 2025-07-04 16:32:54 +08:00
  • 1e3e3add3d fix(docs): fix the broken link in docs/references/production_metrics.md (#7741) Yuchen Cheng 2025-07-04 14:46:07 +08:00
  • 8c298031d5 refactor llama4 dp attention logic (#7729) Yi Zhang 2025-07-04 13:48:11 +08:00
  • 4de0395343 Add V2-lite model test (#7390) YanbingJiang 2025-07-04 13:25:50 +08:00
  • 8b1942c6cc Remove type conversion and fix id map in topk (#7759) Ke Bao 2025-07-04 09:13:32 +08:00
  • 489934be0a fuse renormal into moe topk softmax kernel python code (#7751) Yi Zhang 2025-07-04 07:22:14 +08:00
  • 43f93f632c fix CI: update native api ipynb (#7754) Xinyuan Tong 2025-07-03 15:25:00 -07:00
  • aca1101a13 chore: bump sgl-kernel 0.2.2 (#7755) Yineng Zhang 2025-07-03 12:49:10 -07:00
  • 2998c4bdf4 [optimize] fuse renormalize into moe_topk_softmax (#7744) Yi Zhang 2025-07-04 03:42:44 +08:00
  • 6840a7bbb2 [fix] put cpu in the first priority in get_device() (#7752) JieXin Liang 2025-07-04 02:49:32 +08:00
  • c01a1df588 [Bug] add flashinfer bool check for fusedmoe in Qwen moe models (#7723) yilian49 2025-07-03 11:32:11 -07:00
  • 0099172327 feat: use D2D instead of H2H in pp (#7673) TianyuZhang1214 2025-07-04 01:58:50 +08:00
  • 264dc6e744 [optimize] add two stream norm for qwen3 (#7740) Yi Zhang 2025-07-04 00:59:17 +08:00
  • 646cef2e2e support qwen3 dense model dp attention (#7681) Yi Zhang 2025-07-04 00:58:20 +08:00
  • 1dce6c480f [CPU] support the case where num_attention_heads or intermediate_size is not divisible by the TP size (#6771) Chunyuan WU 2025-07-04 00:51:38 +08:00
  • 9fcc9a80e7 [CPU] refine CPU integration code (#7647) Chunyuan WU 2025-07-04 00:51:09 +08:00
  • ac49dac009 [fix] fix dsv3_router_gemm filter (#7750) JieXin Liang 2025-07-04 00:25:32 +08:00
  • 1e0e549766 Ascend attention backend(PA&MLA) (#7722) ronnie_zheng 2025-07-03 19:23:19 +03:00
  • b58226510f fix dsv3 fused proj check (#7738) AniZpZ 2025-07-03 16:52:44 +08:00
  • 2c4feaf308 Add CUTLASS FP8 Blockscale MoE kernel for Hopper architecture (#7278) ayrnb 2025-07-03 14:27:03 +08:00
  • 2ff572e28c [CI][Router] Fix bench_one_batch_server for pd router test (#7731) Shangming Cai 2025-07-03 14:18:24 +08:00
  • 84f2e4a0f8 fix awq and dsv3 fused gemm compatible (#7735) AniZpZ 2025-07-03 13:56:57 +08:00
  • 8f844db699 [CPU] fix all_reduce and all_gather (#6770) Chunyuan WU 2025-07-03 13:39:45 +08:00
  • 36cc3ffdc7 [CPU] [sgl-kernel] set dispatch key of initialize to CatchAll (#7734) Chunyuan WU 2025-07-03 13:39:24 +08:00
  • 1bebd3154e Fix num_tokens_pre_allocated in disaggregation log (#7714) Ziming Huang 2025-07-03 13:31:49 +08:00
  • d3c275b117 Support updating weights at once by stopping all requests (#6698) Albert 2025-07-03 13:26:06 +08:00
  • b044400dd3 Support non-contiguous query input for extend/decode attention (#7462) YanbingJiang 2025-07-03 10:59:45 +08:00
  • 40e5cb7a9c [CPU] Bind threads and numa node for each TP rank (#6549) Chunyuan WU 2025-07-03 10:57:59 +08:00
  • 8e64140e35 [b200] support trt-llm allreduce fuse rms_norm_add kernel (#7621) Xiaoyu Zhang 2025-07-03 10:36:20 +08:00
  • 82f021e22e [router] add --log-level to sgl-router (#6512) Zilin Zhu 2025-07-03 10:33:04 +08:00
  • 0626f678de [RL] support update_weights_from_distributed with different group and multiple weights (#7292) Zilin Zhu 2025-07-03 10:29:11 +08:00
  • 09e699bba4 [RL] add --skip-warmup (#7416) Zilin Zhu 2025-07-03 09:50:43 +08:00
  • b116b21a46 [AMD] Temporarily disable test_no_overlap_scheduler and test_vision_chunked_prefill (#7717) Hubert Lu 2025-07-02 12:39:18 -07:00
  • 88f484ce4c Apply dsv3 router gemm kernel for deepseek-r1 fp4 (#7677) Baizhou Zhang 2025-07-02 12:30:18 -07:00
  • 8e03b641ba [1/n] apply wna16marlin kernel in moe weight only quantization (#7683) AniZpZ 2025-07-02 14:21:25 +08:00
  • b3fa5dc3c8 Fix GPTQMarlinMoE (#7697) Kyungmin Lee 2025-07-02 14:34:43 +09:00
  • 00aec6ad6c Apply dsv3_fused_a_gemm kernel (#7635) Ke Bao 2025-07-02 13:32:05 +08:00
  • 1a08358aed Improve error handling for requests with unloaded LoRA path(s) (#7642) Lifu Huang 2025-07-01 20:05:34 -07:00
  • f18a8fddd4 chore: upgrade flashinfer v0.2.7.post1 (#7698) Yineng Zhang 2025-07-01 14:05:57 -07:00
  • a7efbb2757 fix(model loader): use safe_open to prevent file handle leaks. (#7684) Simon_CQK 2025-07-02 04:18:35 +08:00
  • 93b6785d78 add description for llama4 eagle3 (#7688) Yi Zhang 2025-07-01 16:19:19 +08:00
  • f9eb04ddb2 upgrade sgl kernel to 0.2.1 for main (#7676) Zhiqiang Xie 2025-07-01 00:00:13 -07:00
  • 3a911b854d Refactor mm processors and Enable mixed modality processing (#7629) Xinyuan Tong 2025-06-30 23:14:48 -07:00
  • 886d344964 support llama4 eagle3 (#6985) lukec 2025-07-01 13:34:10 +08:00
  • 637bfee448 chore: bump sgl-kernel v0.2.1 (#7675) Yineng Zhang 2025-06-30 22:12:33 -07:00
  • 6005eceee3 [CPU] remove process_group from inputs of shm_allreduce and shm_allgather (#7486) Chunyuan WU 2025-07-01 12:54:11 +08:00
  • ff2e9c9479 Add small requirements for benchmark/parse_result tools (#7671) Xiaoyu Zhang 2025-07-01 12:52:20 +08:00
  • 3e34e9004f Fix: sync prepare_fp8_layer_for_marlin with latest vllm changes (#7648) narutolhy 2025-06-30 21:51:01 -07:00
  • 7349717e4b [doc] update lws doc for pd (#7318) ybyang 2025-07-01 10:39:04 +08:00
  • 392e441ad1 chore: upgrade flashinfer v0.2.7 jit (#7663) Yineng Zhang 2025-06-30 13:26:26 -07:00
  • 7248272ccc Add dsv3 router gemm kernel (#7627) Baizhou Zhang 2025-06-29 23:31:55 -07:00
  • 22352d47a9 Improve streaming, log_level, memory report, weight loading, and benchmark script (#7632) Lianmin Zheng 2025-06-29 23:16:19 -07:00
  • c5131f7a2f [CPU] add c++ kernel to bind CPU cores and memory node (#7524) Chunyuan WU 2025-06-30 10:45:25 +08:00
  • 78700893ee [EAGLE] remove a wrong adjustment for page_size > 1 & topk > 1 in server_args.py (#7643) Lianmin Zheng 2025-06-29 19:25:28 -07:00
  • 663c04f76e Update CODEOWNERS (#7640) Lianmin Zheng 2025-06-29 16:58:43 -07:00
  • 3b3f1e3aeb [AMD] Add unit-test-sgl-kernel-amd to AMD CI (#7539) Hubert Lu 2025-06-29 15:50:09 -07:00
  • b691dcc490 [misc] reduce weird rope_scaling_factor warning (#7176) JieXin Liang 2025-06-30 06:42:45 +08:00
  • 0c9c6c75a8 Move files related to EPLB (#7580) fzyzcjy 2025-06-30 06:39:38 +08:00
  • e3f9b54819 [bugfix] fix runtime dropping panic in editable (#7628) Simo Lin 2025-06-29 15:38:28 -07:00
  • b3cff3651e Fix sgl-router startup crash (#7619) finetune 2025-06-29 23:41:34 +02:00
  • 8f335b5bd6 Fix stream reasoning parser and Adds Kimi reasoning parser (#7432) Xinyuan Tong 2025-06-29 14:39:05 -07:00
  • b2264076dc Add @mickqian as the CODEOWNERS of multimodal (#7636) Lianmin Zheng 2025-06-29 09:27:33 -07:00
  • 04b35190e2 Add dsv3 fused a gemm to sgl-kernel (#7630) Ke Bao 2025-06-29 17:52:24 +08:00
  • 071a1f51ae [Minor] clean up multimodal processor and tokenizer manager (#7624) Lianmin Zheng 2025-06-29 02:50:14 -07:00
  • 7c0db3a6c5 [bugfix] Remove PR comment posting from Rust benchmark workflow (#7625) Simo Lin 2025-06-28 22:10:01 -07:00
  • c45e49d817 oai: Adds support for OpenAI chat completions API in bench_serving (#7036) Xinyuan Tong 2025-06-28 15:59:20 -07:00
  • d80539291b docs: add gb200 nvl72 and a16z grant (#7620) Yineng Zhang 2025-06-28 02:08:09 -07:00
  • 00c7b1ad07 Let EP prefill support new DeepGEMM (#7310) fzyzcjy 2025-06-28 16:45:30 +08:00
  • 82eccae44e Let ep_scatter support arbitrary strides / ue8m0 format (#7309) fzyzcjy 2025-06-28 16:38:33 +08:00
  • a8c10aeeee fix unit tests (#7618) Yineng Zhang 2025-06-28 00:32:41 -07:00
  • eb429b88a4 [PD] Respect sampling_params.max_new_tokens when PD disaggregation is activated (#7598) Shangming Cai 2025-06-28 13:22:01 +08:00
  • 49538d111b Support dynamic LoRA loading / unloading in engine/server API (#7446) Lifu Huang 2025-06-27 21:00:27 -07:00
  • cfe2edac38 [BUG] fix local_rank in initialize_dp_attention (#7584) Sheng Qi 2025-06-28 11:01:01 +08:00
  • 2373faa317 Fix flakiness in LoRA batch test. (#7552) Lifu Huang 2025-06-27 19:51:43 -07:00
  • 9efb2993da Tiny add logs for expert location updater (#7308) fzyzcjy 2025-06-28 10:12:33 +08:00
  • a5317b2fd3 [CPU] add optimizations for INT8 and FP8 DeepSeek (#6769) Chunyuan WU 2025-06-28 10:04:29 +08:00
  • eb6c2c1663 Hybrid kv cache for LLaMA4 (#6563) tarinkk 2025-06-27 21:58:55 -04:00
  • 357921aa51 Fix: Minicpm (#7612) Xinyuan Tong 2025-06-27 17:32:29 -07:00
  • c071198c1d [router] add centralized configuration module for sgl-router (#7588) Simo Lin 2025-06-27 15:42:02 -07:00
  • d7374d7467 Fix broken CI TestVILAServer (#7610) Lifu Huang 2025-06-27 15:01:03 -07:00
  • ce3a3e8783 Move multimodal processors into a separate folder (#7581) Lianmin Zheng 2025-06-27 11:58:24 -07:00
  • 41650b0d70 feat: support compatibility between MTP and two-batch-overlap (#7225) Qiaolin Yu 2025-06-27 01:10:27 -07:00
  • 1b95162008 Updates transformers and timm dependencies (#7577) Xinyuan Tong 2025-06-27 00:30:17 -07:00