Commit Graph
100 Commits
Author SHA1 Message Date
Baizhou Zhang c36a10aabb Tiny update pull-requests permission of release-branch-cut.yml (#19121) 2026-02-21 20:14:31 +08:00
Baizhou Zhang 9a32f8ccb9 [CI] Move test_load_lora_from_tensor test to H100 (#18797) 2026-02-13 21:28:00 +08:00
Baizhou Zhang 947927bdb5 [V3.2] Change default CP token split method to --round-robin-split (#18613) 2026-02-11 20:14:35 +08:00
Baizhou Zhang 2d38b8aca0 Revert "[sgl-kernel] upgrade deepgemm" (#18562) 2026-02-11 01:17:40 +08:00
Baizhou Zhang 615a02dcd4 Revert "optimize get_topk_ragged by fusing get k and k_scale triton kernel" (#18471) 2026-02-09 16:37:19 +08:00
Baizhou Zhang eb4cf1dfc4 [CI] Skip some flaky subtests for test_multi_lora_backend.py (#18408) 2026-02-07 19:06:53 +08:00
Baizhou Zhang 9fbec79906 Revert "[Build] Enable full kernel in aarch64 wheel" (#18385) 2026-02-07 09:19:07 +08:00
Baizhou Zhang f2e0048d06 Add CI permission for Shunkangz, dongjiyingdjy, samuellees (#18377) 2026-02-07 01:19:02 +08:00
Baizhou Zhang d279520ba5 [DeepGemm] Add a flag for fast warmup (#18111) 2026-02-04 14:12:13 +08:00
Baizhou Zhang c7d53fa26a Set torch url index in pyproject.toml (#16802) 2026-02-01 13:23:52 +08:00
Baizhou Zhang 1d942e4eef [DeepSeek] Update tests and document for DeepSeek V3.2 NVFP4 checkpoint (#17657) 2026-01-27 22:10:57 +08:00
Baizhou Zhang 832c756549 [Doc] Tiny update description on torch compile (#17819) 2026-01-27 18:59:04 +08:00
Baizhou Zhang 0dfe46dafb [Docker] Install cudnn==9.16 for cuda 13 image to avoid check error (#17668) 2026-01-24 11:27:03 +08:00
Baizhou Zhang 283a2daeaa [hotfix] Reenable all reduce fusion on sm100 (#17591) 2026-01-22 23:36:38 +08:00
Baizhou Zhang 8dae6ec03c Add xyjixyjixyji to CI_Permission (#17559) 2026-01-21 23:27:25 -08:00
Baizhou Zhang e2d33531f3 [Kernel] Little refactor of flashinfer allreduce norm fusion (#17474) 2026-01-22 13:31:57 +08:00
Baizhou Zhangandiforgetmyname fafa171529 [hotfix] Fixes on cuda 13 docker image (#17541)
Co-authored-by: iforgetmyname <iforgetmyname@users.noreply.github>
2026-01-22 12:29:55 +08:00
Baizhou Zhang 3373545b9f [HotFix]Fix dtype mismatch in nsa indexer on AMD device (#17518) 2026-01-22 12:17:02 +08:00
Baizhou Zhang 8251a74d5f [Tiny] Backward compatibility for fp4 gemm flags (#17466) 2026-01-21 14:34:40 +08:00
Baizhou Zhang a54d75bf2e [Fix] Set fa3 as default MHA backend on Hopper (#17425) 2026-01-21 13:54:09 +08:00
Baizhou Zhang c3f9c30f99 [Minor] Change lora_target_modules to "all" in CI tests (#17386) 2026-01-21 11:46:36 +08:00
Baizhou Zhang 6ea491e439 Overlap shared experts with deepep dispatch for single batch overlap on Blackwell (#17289) 2026-01-21 02:56:55 +08:00
Baizhou Zhang 55c616427d Add flag that enables NCCL mlp sync batch for overlap scheduler (#17288) 2026-01-20 23:06:55 +08:00
Baizhou Zhang ea879c7739 [Minor] Correct sglang version when installing from source (#17315) 2026-01-18 19:36:16 -08:00
Baizhou Zhang 8b9e9357fe [2/n] deepseek_v2.py Refactor: Migrate MHA forward method in deepseek_v2.py (#16817) 2026-01-17 09:36:25 +08:00
Baizhou Zhang a04675892e Update flashinfer to 0.6.1 (#15551) 2026-01-17 00:48:30 +08:00
Baizhou Zhang 8b99af9af8 [Doc] Tiny update Cuda 13 environment instructions (#17174) 2026-01-16 06:12:26 +08:00
Baizhou Zhang f9fc50acd6 [Tiny] Rename test_sparse_flash_attn.py to fix CI (#16895) 2026-01-11 18:18:29 +08:00
Baizhou Zhang 8b5d426340 [CI]Move fa4 e2e test to 4-gpu-b200 runner (#16889) 2026-01-11 15:53:38 +08:00
Baizhou Zhang 9fd2358cc2 Update Cutedsl version and pin cuda-python version (#16838) 2026-01-10 17:08:43 +08:00
Baizhou Zhang 7f393d9512 [Docker] Add nightly dev docker for Cuda 13 (#16862) 2026-01-10 14:56:53 +08:00
Baizhou Zhang 94fc26aad8 [Doc]Update note for Cuda 13 container usage (#16805) 2026-01-10 14:03:19 +08:00
Baizhou Zhang 38dc5839dd [1/n]deepseek_v2.py Refactor: attention backend handlers and forward method definition (#16306) 2026-01-08 09:22:31 +08:00
Baizhou Zhang 153c69f63d [CI] Enable dpsk v31 test on nightly H200 (#16660) 2026-01-07 23:21:19 +08:00
Baizhou Zhang 7d757d6f17 Clean Some Environment Variables for DeepSeek V32 (#15938) 2026-01-07 14:00:16 +08:00
Baizhou Zhang 6ffe1fc02f [Fix]Pin mooncake version to 0.3.7.post2 in grace blackwell (#16502) 2026-01-06 14:11:40 +08:00
Baizhou Zhang bb23a8fe77 [Tiny]Remove progress bar for fp8 ue8m0 quant when unneeded (#16177) 2026-01-03 10:54:12 +08:00
Baizhou Zhang f07e76b229 Multiple refactors of DeepSeek V32 and context parallel (#16305) 2026-01-03 02:21:22 +08:00
Baizhou Zhang 70a769bc56 Fix NPU docker release workflow (#16253) 2026-01-01 10:53:45 +08:00
Baizhou Zhang e47afa0237 [DP]Fix sync bubble in adjust_num_token_non_padded_for_attn_tp (#16178) 2025-12-31 17:58:18 +08:00
Baizhou Zhang f35b5da521 [CI] Append test variant name to markdown report header in nightly test (#16166) 2025-12-31 00:09:24 +08:00
Baizhou ZhangandKangyan Zhou 98225be6e5 [CI] Fixing release with cut branch workflow (#16153)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
2025-12-30 17:18:02 +08:00
Baizhou Zhang 208e6a9dac [Doc]Update MTP moe backends for EP document (#16013) 2025-12-28 19:36:23 +08:00
Baizhou Zhang 656f4d69a1 Refactor fp8 nextn layer for DeepSeek nvfp4 checkpoint (#15353) 2025-12-28 11:57:09 +08:00
Baizhou Zhang 468931b572 [Tiny]Move deepseek fp4 cutlass moe test to per-commit test (#15565) 2025-12-21 18:08:07 -08:00
Baizhou Zhang 2b0ddf89f5 [Tiny]Add warning for deepgemm on Blackwell (#15352) 2025-12-18 13:45:26 -08:00
Baizhou Zhang 8451e22758 [DeepSeek-V32]Update nightly performance benchmark (#15308) 2025-12-17 11:25:31 -08:00
Baizhou Zhang d92c1f8cbd Fix lora doc (#15282) 2025-12-16 15:53:38 -08:00
Baizhou Zhang 28a19e494b Fix lint (#15281) 2025-12-16 13:36:01 -08:00
Baizhou Zhang 0261c4aff7 [misc] Upgrade cutedsl to 4.3.1 (#14857) 2025-12-16 12:11:56 -08:00
Baizhou Zhang c843419562 Remove duplicate bs=1 in nightly benchmark (#15162) 2025-12-15 22:09:22 -08:00
Baizhou Zhang ab3ffd1c8e Add nightly accuracy test for DeepSeek V3.2 (#14935) 2025-12-13 12:11:16 -08:00
Baizhou Zhang 8698867479 [CI]Add gb200 runner back (#15024) 2025-12-12 20:19:34 -08:00
Baizhou Zhang 7dcad45cad [CI] Temp disable gb200 test (#14865) 2025-12-10 19:05:16 -08:00
Baizhou Zhang e5201bda34 [CI] Unblock gb200 cutedsl test (#14469) 2025-12-08 17:58:25 -08:00
Baizhou Zhang 6799847ebf [CI]Unblock and split spec v2+dp test (#14551) 2025-12-07 17:39:25 -08:00
Baizhou Zhang 673c11ba73 [Minor] Temporarily skipping deepep large mtp test (#14586) 2025-12-07 13:59:16 -08:00
Baizhou Zhang 9dfa01a435 [Misc]Register and refactor some environs for dpsk-fp4 and DeepEp (#14538) 2025-12-06 12:29:16 -08:00
Baizhou Zhang bc388471d2 [1/n] Fix hanging during DeepGemm Warmup (#14493) 2025-12-06 10:44:02 -08:00
Baizhou Zhang 42fcf5438f Revert "tiny remove deprecated endpoint call" (#14533) 2025-12-05 23:48:54 -08:00
Baizhou Zhang 80a575e4e8 Add YAMY1234 to CI Permission (#14475) 2025-12-04 21:25:49 -08:00
Baizhou ZhangandXinyuan Tong 7e78825d5a [Tiny]Small fixes in deepseek v32 doc (#14372)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2025-12-03 11:35:40 -08:00
Baizhou Zhang 4bcc5879af [Doc] Fix DeepSeek V32 Doc (#14336) 2025-12-02 21:06:55 -08:00
Baizhou Zhang 922054079c [Doc] Update DeepSeek-V3.2 document (#14321) 2025-12-02 18:19:39 -08:00
Baizhou Zhang 03888b9de5 [Minor] Upgrade cutedsl version in Dockerfile (#13968) 2025-12-01 17:15:26 -08:00
Baizhou Zhang eb5008846a [CI] Fix test_deepep_large.py (#14247) 2025-12-01 15:18:48 -08:00
Baizhou Zhang f1115cf58d Revert "[Minor]Raise Error when deepep num dispatch token per rank is smaller than cuda graph bs" (#14171) 2025-11-30 12:49:46 -08:00
Baizhou Zhang 7b03cc6482 [Minor]Raise Error when deepep num dispatch token per rank is smaller than cuda graph bs (#14065) 2025-11-30 10:11:42 -08:00
Baizhou Zhang 051ad83347 [chore] Arrange NV packages in Dockerfile (#13749) 2025-11-27 18:08:27 -05:00
Baizhou Zhang 7ab548ef64 [2/2] Refactor DeepGeem requant for FP8 FusedMoE on Blackwell (#13960) 2025-11-27 09:00:26 -05:00
Baizhou Zhang 8a9b8b8457 Revert "Fix nightly test failures: NSA indexer dtype and CPP radix cache init" (#14015) 2025-11-26 10:45:23 -08:00
Baizhou Zhang 873382a910 [Tiny]Upgrade README for sgl-kernel (#13945) 2025-11-25 16:46:30 -08:00
Baizhou Zhang 808b6dfdea [Minor] Fix lint (#13938) 2025-11-25 10:57:23 -08:00
Baizhou Zhang 04b52fa8d6 [chore]Upgrade flashinfer to 0.5.3 (#13751) 2025-11-23 23:38:36 -08:00
Baizhou Zhangandfy1214 4683e244fe [1/2] Refactor DeepGeem requant for FP8 Linear on Blackwell (#13601)
Co-authored-by: fy1214
2025-11-23 16:07:56 -08:00
Baizhou Zhang c9bd1aca32 [CI] Tiny refactoring sgl-kernel tests (#13813) 2025-11-23 12:45:17 -08:00
Baizhou Zhang 8bfce9b08d [Tiny] Renaming environ for NVFP4 dispatch (#13756) 2025-11-22 00:05:20 -08:00
Baizhou Zhang 9f59194f29 [Fix] Fix DeepSeek V3 MTP on B200 (#13548) 2025-11-18 16:30:26 -08:00
Baizhou Zhang 10969ae4be [chore] Disable ccache for sgl-kernel release (#13541) 2025-11-18 14:28:58 -08:00
Baizhou Zhang 85ae508e8b Add bfloat16 tuned fused moe config for Dpsk-MTP layer on B200 (#13455) 2025-11-17 17:44:31 -08:00
Baizhou Zhang d64dd3e18e [Tiny]Fix 1-gpu nightly test bugs (#13389) 2025-11-16 15:54:17 -08:00
Baizhou Zhang 3ccd7fa669 [CI] Fix B200 CI (#13387) 2025-11-16 15:13:57 -08:00
Baizhou Zhang 10285ec204 [Misc]Add date to cu13 dev image tag (#13316) 2025-11-14 19:42:50 -08:00
Baizhou Zhang 8ece99a9dc [CI] Update job dependency and move dpsk v3.2 tests to 8-gpu suite (#12942) 2025-11-11 23:16:47 -08:00
Baizhou Zhang 99e25805f5 [Fix] Fix nan error for large scale ep (#12866) 2025-11-11 14:44:57 -08:00
Baizhou Zhang 5f02b918ec [Fix] Fix trtllm-mla backend when chunked prefix cache is disabled (#12361) 2025-11-08 15:10:25 -08:00
Baizhou Zhang e039ff382c [CI] Fix huggingface access for test_flash_attention_4.py (#12846) 2025-11-07 20:07:06 -08:00
Baizhou Zhang bb6a21cd99 [Fix]Tiny fix in Dockerfile (#12748) 2025-11-05 21:04:33 -08:00
Baizhou Zhang 3c0a6df82d [chore] Fix triton installation for cu13 image (#12742) 2025-11-05 20:20:18 -08:00
Baizhou Zhang 9a954982de [chore] SGLang tag management in Dockerfile (#12734) 2025-11-05 19:03:26 -08:00
Baizhou Zhang 9ec6031d70 [chore]Remove dockerfile from target file of bump kernel version (#12728) 2025-11-05 18:01:19 -08:00
Baizhou Zhang 7c45b8b4bb [CI] Fix qwen3-vl lora nightly ci (#12708) 2025-11-05 11:00:13 -08:00
Baizhou Zhang d22d044734 Revert "Enable memory saver for hybrid model" (#12648) 2025-11-04 16:22:06 -08:00
Baizhou Zhang 42889acbd0 [hotfix] Fix deepep w4a8 bug (#12642) 2025-11-04 13:55:59 -08:00
Baizhou Zhang 15efbcb4e7 [chore] Fix update_kernel_whl_index script for multiple cuda version (#12519) 2025-11-03 16:34:14 -08:00
Baizhou Zhang 6e29446e45 [hotfix] Remove flashinfer-jit-cache from pyproject (#12530) 2025-11-02 22:11:05 -08:00
Baizhou Zhang 9a512cf95b [CI] Move some Lora/Deterministic CI tests to nightly (#12507) 2025-11-01 19:54:22 -07:00
Baizhou Zhang 2b7bf11bd2 [Hotfix] Remove extra comment in sgl-kernel README (#12500) 2025-11-01 12:22:45 -07:00
Baizhou Zhang 566ade0388 [CI] Build aarch64 kernels for sgl-kernel test (#12480) 2025-11-01 11:55:42 -07:00
Baizhou Zhang 5f98b7fe61 [CI] Fix kernel installation on aarch runners (#12475) 2025-10-31 14:25:27 -07:00