Commit Graph

  • 89caf7a3c6 [bugfix] Apply routed scaling factor to cutlass_fused_experts_fp8 (#8688) Trevor Morris 2025-08-01 19:00:24 -07:00
  • b27b11919c chore(gb200): update dockerfile to handle fp4 disaggregation (#8694) ishandhanani 2025-08-01 18:58:00 -07:00
  • f642524fd9 [1/2] sgl-kernel: Fuse routed scaling factor into select_experts (#8364) Trevor Morris 2025-08-01 18:14:24 -07:00
  • 82e6c3a65a Add support for NCCL symmetric memory for TP allreduces (#8238) Nicolas Castet 2025-08-01 18:30:55 -05:00
  • b89d37cb11 [bugfix] Add 'disaggregation_mode' parameter to warmup function when compile deep_gemm manually (#8618) Baron Liu 2025-08-02 07:02:53 +08:00
  • 5deab1283a upgrade xgrammar 0.1.22 (#8522) Swipe4057 2025-08-02 01:59:15 +03:00
  • d1c4d51c08 bugfix(hicache): Fix 'MooncakeStore' not defined error. (#8668) hzh0425 2025-08-02 06:58:17 +08:00
  • 1fe691a429 Fix FP8 block quantization when N or K is not multiples of 128 (#8648) YanbingJiang 2025-08-02 06:57:19 +08:00
  • e252192679 Fix deepgemm masked grouped gemm jit compile (#8679) Ke Bao 2025-08-02 06:37:59 +08:00
  • 07e46ecaad Update CODEOWNERS (#8686) Lianmin Zheng 2025-08-01 15:09:44 -07:00
  • ab9b893e61 [bug] limit bootstrap room to to [0, 2^63 - 1] (#8684) Simo Lin 2025-08-01 14:41:01 -07:00
  • 6a7528e623 [bugfix] Fix page size for create_flashmla_kv_indices_triton() for cutlass mla (#8685) Trevor Morris 2025-08-01 14:28:04 -07:00
  • 2ae95d17e8 Disable tp for shared experts under expert parallelism for GLM4.5 model (#8647) (#8647) Minglei Zhu 2025-08-01 12:02:35 -07:00
  • 2d401bd99d [fix] fix pd disagg error of vlms (#8094) 萝卜菜 2025-08-02 02:16:29 +08:00
  • b17c5b0118 fix arg typo for --disaggregation-transfer-backend (#8664) Zac 2025-08-02 01:00:47 +08:00
  • db7343c992 fix per token cuda kernel hidden dim cannot divide by 16 (#8543) Stefan He 2025-08-01 09:27:18 -07:00
  • 533cb5b274 [DOC]Update sgl-kernel README (#8665) Hongbo Xu 2025-08-01 22:59:27 +08:00
  • 6bdd27861b [Kimi K2] dsv3_router_gemm supports NUM_EXPERTS == 384 (#8013) Peter Pan 2025-08-01 22:01:24 +08:00
  • 46e9d1c7c1 Increase tolerance to address CI failures (#8643) Lifu Huang 2025-08-01 02:32:10 -07:00
  • 6c88f6c8d9 [5/N] MoE Refactor: Update MoE parallelism arguments (#8658) Cheng Wan 2025-08-01 01:20:03 -07:00
  • c8d3a402c1 Bug: apply final_hidden_states*=self.routed_scaling_factor at MoE lay… (#8511) Binyao Jiang 2025-08-01 00:07:41 -07:00
  • 7e831efee8 Fix chat template handling for OpenAI serving (#8635) Xinyuan Tong 2025-07-31 21:49:45 -07:00
  • 20b5563eda Add hf3fs_utils.cpp to package-data (#8653) pansicheng 2025-08-01 12:41:09 +08:00
  • 33f0de337d chore: bump v0.4.10.post1 (#8652) Ke Bao 2025-08-01 12:07:30 +08:00
  • e7e5a3050a Update batch size limitation of dsv3_router_gemm kernel to 16 (#8051) Baizhou Zhang 2025-07-31 20:53:31 -07:00
  • dd7ca00601 Interface change for kvcache io to support page first layout (#8318) Zhiqiang Xie 2025-07-31 20:37:49 -07:00
  • 9305ea6c2d HiCache, fixing hash value indexing (#8636) Zhiqiang Xie 2025-07-31 20:29:51 -07:00
  • aa4c66b564 [NVIDIA] Enable Flashinfer MoE blockscale fp8 backend for TP MoE (#8450) Kaixi Hou 2025-07-31 19:56:34 -07:00
  • 39decec10b [router] upgrade router version to 0.1.8 (#8645) Simo Lin 2025-07-31 19:00:23 -07:00
  • f6f46f4629 [router] add basic usage doc (#8640) Simo Lin 2025-07-31 18:11:48 -07:00
  • 2886e23dbd [bugfix] fix router python parser for pd urls (#8644) Simo Lin 2025-07-31 18:09:31 -07:00
  • 99795d61e6 [Bugfix] fix w8a8_int8 load issue (#8308) Even Zhou 2025-08-01 08:30:16 +08:00
  • fe5086fd8b chore: speedup NPU CI by cache (#8270) li chaoran 2025-08-01 08:29:50 +08:00
  • 04913430c6 Feature/modelscope model download (#8083) yrk111222 2025-08-01 08:29:31 +08:00
  • 0ad098b494 Revert "Fix nan value generated after custom all reduce (#8532)" (#8642) Yineng Zhang 2025-07-31 17:26:49 -07:00
  • 4a6e7a66a0 Fix nan value generated after custom all reduce (#8532) kk 2025-08-01 07:15:43 +08:00
  • 4b04998d38 TRTLLM Gen MLA Decode Kernel Integration (same as #7938) (#8632) Faraz 2025-07-31 19:03:40 -04:00
  • 3dde86194a Conditionally import HiCacheHF3FS (#8598) pansicheng 2025-08-01 05:59:29 +08:00
  • b7170cc820 [bugfix] Fix flashinfer cutlass EP moe after MoE refactor (#8630) Trevor Morris 2025-07-31 13:57:08 -07:00
  • 5c14515fec [bug] remove pdlb from minilb since its no longer available (#8634) Simo Lin 2025-07-31 13:54:02 -07:00
  • 2cd2e27f80 SGLang HiCache NIXL Connector (#8488) Vishwanath Venkatesan 2025-07-31 15:09:42 -05:00
  • 743638bc03 misc: Remove debug print to logger.info (#8633) Chang Su 2025-07-31 12:56:52 -07:00
  • 061c8959ff Fix typos in py_test/test_launch_server.py (#6227) Michael Yao 2025-08-01 03:48:47 +08:00
  • 4acf690206 [Optimization][Perf] Disable the GC during CUDA graph capture to speed up by up to 3x (#8577) Brayden Zhong 2025-07-31 14:31:21 -04:00
  • aee0ef52f5 [router] update router pypi version (#8628) Simo Lin 2025-07-31 11:24:12 -07:00
  • ae807774f5 [ci] fix genai-bench execution cmd (#8629) Simo Lin 2025-07-31 10:40:54 -07:00
  • 8fbcfd0723 Update step3v default config (#8626) Ke Bao 2025-08-01 00:49:26 +08:00
  • 3c307dc057 Fix hf3fs_fuse import error (#8623) Ke Bao 2025-07-31 22:42:31 +08:00
  • 5d15fb8c9d [bugifx] QWen-1M context support[2/3] using current cuda stream in the DCA's kernel for bugfix. (#8611) Tao He 2025-07-31 22:41:39 +08:00
  • 016fd25127 [PD] Use batch transfer for rdma transport and add notes for mnnvl usage (#8595) Shangming Cai 2025-07-31 21:29:34 +08:00
  • 023288645b chore: bump v0.4.10 (#8608) Yineng Zhang 2025-07-31 05:50:17 -07:00
  • 7a1f7fc504 [Feature] Hybrid EP and TP (#8590) Cheng Wan 2025-07-31 02:53:25 -07:00
  • 51c38163c1 model: support Step3V (#8583) Chang Su 2025-07-31 02:41:00 -07:00
  • 09f1a247ce fix: fork should not run pypi router (#8604) yihong 2025-07-31 17:37:13 +08:00
  • 32fa1e9cc2 [4/N] MoE Refactor: Unified Triton Kernel for FusedMoE and EPMoE (#8515) Cheng Wan 2025-07-31 02:34:02 -07:00
  • e7dc163f57 add SVG logo (#8603) Liangsheng Yin 2025-07-31 15:56:26 +08:00
  • e179e0b797 update sgl-kernel for EP: python part (#8550) Cheng Wan 2025-07-31 00:14:39 -07:00
  • d904959233 Support l3 cache (mooncake store) for hiradix cache (#7211) huangtingwei 2025-07-31 14:15:51 +08:00
  • 26c8a310bd fix incorrect increase of hit count (#8533) huangtingwei 2025-07-31 14:02:42 +08:00
  • 5963e50503 [bugfix] Fix 2 minor bugs in the hicache storage layer (#8404) yi wang 2025-07-31 13:47:14 +08:00
  • 43118f5f2a chore: bump sgl-kernel v0.2.8 (#8599) Yineng Zhang 2025-07-30 22:23:52 -07:00
  • a5f5ab4030 update sgl-kernel for EP: kernel part (#8514) Cheng Wan 2025-07-30 22:19:55 -07:00
  • 59aab76f0a Bug: Fix google gemma3n-mm audio input not working bug (#8365) Binyao Jiang 2025-07-30 21:23:09 -07:00
  • 659bfd1023 Add GKE's default CUDA runtime lib location to PATH and LD_LIBRARY_PATH. (#8544) Charles Chen 2025-07-30 20:28:07 -07:00
  • 67e53b16f5 Bump transfomers to 4.54.1 to fix Gemma cache issue. (#8541) Lifu Huang 2025-07-30 19:50:54 -07:00
  • 9b9e82539b [Fix]Fix index oob in get_group_gemm_starts kernel. (#8564) Qi Yuhang 2025-07-31 10:49:35 +08:00
  • 66a398f49d [router] migrate router from actix to axum (#8479) Simo Lin 2025-07-30 17:47:19 -07:00
  • 299803343d Add hf3fs support for hicache storage (based on #7704) (#7280) pansicheng 2025-07-31 08:42:41 +08:00
  • a79a5d7012 Revert "Fix the input tools format and history tool_calls in OpenAI API (#6556)" (#8584) Chang Su 2025-07-30 13:12:05 -07:00
  • ec5f944271 [Model] Add support for Arcee Foundational Model (#8154) Adarsh Shirawalmath 2025-07-30 23:15:25 +05:30
  • 3bdcdd134b [Hot-Fix] moe_aligned_block_size CI failed in AMD (#8461) Yuan Luo 2025-07-31 00:28:32 +08:00
  • a730ce8162 [feature] [sgl-router] Add a dp-aware routing strategy (#6869) Rui Chen 2025-07-30 20:58:48 +08:00
  • 55ecdc0a8e Update CODEOWNERS (#8562) Shangming Cai 2025-07-30 16:05:57 +08:00
  • e3f08c77bc Update cutlass_moe.py (#8545) Elfie Guo 2025-07-29 23:46:34 -07:00
  • a9fd80336d [router] allow longer time out for router e2e (#8560) Simo Lin 2025-07-29 23:43:37 -07:00
  • 2fbb754e1d feature(pd-hicache): Prefill instances support reusing the RemoteStorage Cache via HiCache. (#8516) hzh0425 2025-07-30 12:19:25 +08:00
  • a85ebf50b8 feat(hicache): support file backend reading directory config form env. (#8498) hzh0425 2025-07-30 12:18:46 +08:00
  • 9effeb5bdd Support EPLB in FusedMoE (#8448) Cheng Wan 2025-07-29 16:02:41 -07:00
  • 1992ef9ba7 fix: temporarily disable cuda-ipc for mm data tensor (#8431) Mick 2025-07-30 06:42:03 +08:00
  • c0fd77e839 bring back kimi vl ci (#8537) Stefan He 2025-07-29 13:14:18 -07:00
  • a4c3b121d8 Split the scheduler into multiple mixin classes to reduce the file size (#8483) Lianmin Zheng 2025-07-29 12:46:50 -07:00
  • 5973675bc3 Fix moe align kernel test (#8531) Ke Bao 2025-07-30 02:03:02 +08:00
  • 4d16c88b6e Update cutlass_moe.py (#8535) Elfie Guo 2025-07-29 10:49:41 -07:00
  • 7a4309cc8a [sgl-kernel performace] fix fp8 quant kernels dispatch __nv_fp8_e4m3 bug to improve performance 10%-20% (#8499) Xiaoyu Zhang 2025-07-29 23:31:54 +08:00
  • 813670660c Update README.md (#8528) Lianmin Zheng 2025-07-29 04:09:19 -07:00
  • 263c9236a0 Always trigger pr-test (#8527) Lianmin Zheng 2025-07-29 04:05:19 -07:00
  • 6478831be9 chore: bump v0.4.9.post6 (#8517) Yineng Zhang 2025-07-29 02:30:07 -07:00
  • 2e1d2d7e66 Add PVC and update resource limits in k8s config (#8489) TimWang 2025-07-29 14:15:31 +08:00
  • fb16fbaf52 Fix incorrect KV cache allocation for MTP models. (#8482) Lifu Huang 2025-07-28 22:54:50 -07:00
  • 0ce84c822b Support colocating requests (#7973) fzyzcjy 2025-07-29 13:51:49 +08:00
  • 59d0bf012f Tiny add warnings for DeepEP when it is suboptimal (#8426) fzyzcjy 2025-07-29 13:51:38 +08:00
  • 7df2c0c2db Reduce memory usage for fp4 moe (#8413) fzyzcjy 2025-07-29 13:51:23 +08:00
  • 69712e6f55 Rename the last step in pr-test.yml as pr-test-finish (#8486) Lianmin Zheng 2025-07-28 19:06:13 -07:00
  • 001bffca62 Update CODEOWNERS (#8485) Lianmin Zheng 2025-07-28 17:57:23 -07:00
  • 7c9697178e [CI]Add genai-bench Performance Validation for PD Router (#8477) Keyang Ru 2025-07-28 16:58:23 -07:00
  • 8240a6b013 chore: add glm 4.5 fp8 tp4 config (#8480) Yineng Zhang 2025-07-28 16:14:01 -07:00
  • 3a04aa4be7 chore: add glm4 fp8 tp8 config (#8478) Yineng Zhang 2025-07-28 16:08:53 -07:00
  • bd51694906 Update codeowner (#8476) Lianmin Zheng 2025-07-28 16:03:49 -07:00
  • 74e7e45710 Fix DEEPEP BF16 compatibility for Deepseek Style model like GLM 4.5 (#8469) Stefan He 2025-07-28 14:36:08 -07:00
  • 1466c1b896 feat: support glm4 tuning (#8473) Yineng Zhang 2025-07-28 14:32:58 -07:00