Baizhou Zhang
|
c36a10aabb
|
Tiny update pull-requests permission of release-branch-cut.yml (#19121)
|
2026-02-21 20:14:31 +08:00 |
|
Baizhou Zhang
|
9a32f8ccb9
|
[CI] Move test_load_lora_from_tensor test to H100 (#18797)
|
2026-02-13 21:28:00 +08:00 |
|
Baizhou Zhang
|
947927bdb5
|
[V3.2] Change default CP token split method to --round-robin-split (#18613)
|
2026-02-11 20:14:35 +08:00 |
|
Baizhou Zhang
|
2d38b8aca0
|
Revert "[sgl-kernel] upgrade deepgemm" (#18562)
|
2026-02-11 01:17:40 +08:00 |
|
Baizhou Zhang
|
615a02dcd4
|
Revert "optimize get_topk_ragged by fusing get k and k_scale triton kernel" (#18471)
|
2026-02-09 16:37:19 +08:00 |
|
Baizhou Zhang
|
eb4cf1dfc4
|
[CI] Skip some flaky subtests for test_multi_lora_backend.py (#18408)
|
2026-02-07 19:06:53 +08:00 |
|
Baizhou Zhang
|
9fbec79906
|
Revert "[Build] Enable full kernel in aarch64 wheel" (#18385)
|
2026-02-07 09:19:07 +08:00 |
|
Baizhou Zhang
|
f2e0048d06
|
Add CI permission for Shunkangz, dongjiyingdjy, samuellees (#18377)
|
2026-02-07 01:19:02 +08:00 |
|
Baizhou Zhang
|
d279520ba5
|
[DeepGemm] Add a flag for fast warmup (#18111)
|
2026-02-04 14:12:13 +08:00 |
|
Baizhou Zhang
|
c7d53fa26a
|
Set torch url index in pyproject.toml (#16802)
|
2026-02-01 13:23:52 +08:00 |
|
Baizhou Zhang
|
1d942e4eef
|
[DeepSeek] Update tests and document for DeepSeek V3.2 NVFP4 checkpoint (#17657)
|
2026-01-27 22:10:57 +08:00 |
|
Baizhou Zhang
|
832c756549
|
[Doc] Tiny update description on torch compile (#17819)
|
2026-01-27 18:59:04 +08:00 |
|
Baizhou Zhang
|
0dfe46dafb
|
[Docker] Install cudnn==9.16 for cuda 13 image to avoid check error (#17668)
|
2026-01-24 11:27:03 +08:00 |
|
Baizhou Zhang
|
283a2daeaa
|
[hotfix] Reenable all reduce fusion on sm100 (#17591)
|
2026-01-22 23:36:38 +08:00 |
|
Baizhou Zhang
|
8dae6ec03c
|
Add xyjixyjixyji to CI_Permission (#17559)
|
2026-01-21 23:27:25 -08:00 |
|
Baizhou Zhang
|
e2d33531f3
|
[Kernel] Little refactor of flashinfer allreduce norm fusion (#17474)
|
2026-01-22 13:31:57 +08:00 |
|
 Baizhou Zhangandiforgetmyname
|
fafa171529
|
[hotfix] Fixes on cuda 13 docker image (#17541)
Co-authored-by: iforgetmyname <iforgetmyname@users.noreply.github>
|
2026-01-22 12:29:55 +08:00 |
|
Baizhou Zhang
|
3373545b9f
|
[HotFix]Fix dtype mismatch in nsa indexer on AMD device (#17518)
|
2026-01-22 12:17:02 +08:00 |
|
Baizhou Zhang
|
8251a74d5f
|
[Tiny] Backward compatibility for fp4 gemm flags (#17466)
|
2026-01-21 14:34:40 +08:00 |
|
Baizhou Zhang
|
a54d75bf2e
|
[Fix] Set fa3 as default MHA backend on Hopper (#17425)
|
2026-01-21 13:54:09 +08:00 |
|
Baizhou Zhang
|
c3f9c30f99
|
[Minor] Change lora_target_modules to "all" in CI tests (#17386)
|
2026-01-21 11:46:36 +08:00 |
|
Baizhou Zhang
|
6ea491e439
|
Overlap shared experts with deepep dispatch for single batch overlap on Blackwell (#17289)
|
2026-01-21 02:56:55 +08:00 |
|
Baizhou Zhang
|
55c616427d
|
Add flag that enables NCCL mlp sync batch for overlap scheduler (#17288)
|
2026-01-20 23:06:55 +08:00 |
|
Baizhou Zhang
|
ea879c7739
|
[Minor] Correct sglang version when installing from source (#17315)
|
2026-01-18 19:36:16 -08:00 |
|
Baizhou Zhang
|
8b9e9357fe
|
[2/n] deepseek_v2.py Refactor: Migrate MHA forward method in deepseek_v2.py (#16817)
|
2026-01-17 09:36:25 +08:00 |
|
Baizhou Zhang
|
a04675892e
|
Update flashinfer to 0.6.1 (#15551)
|
2026-01-17 00:48:30 +08:00 |
|
Baizhou Zhang
|
8b99af9af8
|
[Doc] Tiny update Cuda 13 environment instructions (#17174)
|
2026-01-16 06:12:26 +08:00 |
|
Baizhou Zhang
|
f9fc50acd6
|
[Tiny] Rename test_sparse_flash_attn.py to fix CI (#16895)
|
2026-01-11 18:18:29 +08:00 |
|
Baizhou Zhang
|
8b5d426340
|
[CI]Move fa4 e2e test to 4-gpu-b200 runner (#16889)
|
2026-01-11 15:53:38 +08:00 |
|
Baizhou Zhang
|
9fd2358cc2
|
Update Cutedsl version and pin cuda-python version (#16838)
|
2026-01-10 17:08:43 +08:00 |
|
Baizhou Zhang
|
7f393d9512
|
[Docker] Add nightly dev docker for Cuda 13 (#16862)
|
2026-01-10 14:56:53 +08:00 |
|
Baizhou Zhang
|
94fc26aad8
|
[Doc]Update note for Cuda 13 container usage (#16805)
|
2026-01-10 14:03:19 +08:00 |
|
Baizhou Zhang
|
38dc5839dd
|
[1/n]deepseek_v2.py Refactor: attention backend handlers and forward method definition (#16306)
|
2026-01-08 09:22:31 +08:00 |
|
Baizhou Zhang
|
153c69f63d
|
[CI] Enable dpsk v31 test on nightly H200 (#16660)
|
2026-01-07 23:21:19 +08:00 |
|
Baizhou Zhang
|
7d757d6f17
|
Clean Some Environment Variables for DeepSeek V32 (#15938)
|
2026-01-07 14:00:16 +08:00 |
|
Baizhou Zhang
|
6ffe1fc02f
|
[Fix]Pin mooncake version to 0.3.7.post2 in grace blackwell (#16502)
|
2026-01-06 14:11:40 +08:00 |
|
Baizhou Zhang
|
bb23a8fe77
|
[Tiny]Remove progress bar for fp8 ue8m0 quant when unneeded (#16177)
|
2026-01-03 10:54:12 +08:00 |
|
Baizhou Zhang
|
f07e76b229
|
Multiple refactors of DeepSeek V32 and context parallel (#16305)
|
2026-01-03 02:21:22 +08:00 |
|
Baizhou Zhang
|
70a769bc56
|
Fix NPU docker release workflow (#16253)
|
2026-01-01 10:53:45 +08:00 |
|
Baizhou Zhang
|
e47afa0237
|
[DP]Fix sync bubble in adjust_num_token_non_padded_for_attn_tp (#16178)
|
2025-12-31 17:58:18 +08:00 |
|
Baizhou Zhang
|
f35b5da521
|
[CI] Append test variant name to markdown report header in nightly test (#16166)
|
2025-12-31 00:09:24 +08:00 |
|
 Baizhou ZhangandKangyan Zhou
|
98225be6e5
|
[CI] Fixing release with cut branch workflow (#16153)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
|
2025-12-30 17:18:02 +08:00 |
|
Baizhou Zhang
|
208e6a9dac
|
[Doc]Update MTP moe backends for EP document (#16013)
|
2025-12-28 19:36:23 +08:00 |
|
Baizhou Zhang
|
656f4d69a1
|
Refactor fp8 nextn layer for DeepSeek nvfp4 checkpoint (#15353)
|
2025-12-28 11:57:09 +08:00 |
|
Baizhou Zhang
|
468931b572
|
[Tiny]Move deepseek fp4 cutlass moe test to per-commit test (#15565)
|
2025-12-21 18:08:07 -08:00 |
|
Baizhou Zhang
|
2b0ddf89f5
|
[Tiny]Add warning for deepgemm on Blackwell (#15352)
|
2025-12-18 13:45:26 -08:00 |
|
Baizhou Zhang
|
8451e22758
|
[DeepSeek-V32]Update nightly performance benchmark (#15308)
|
2025-12-17 11:25:31 -08:00 |
|
Baizhou Zhang
|
d92c1f8cbd
|
Fix lora doc (#15282)
|
2025-12-16 15:53:38 -08:00 |
|
Baizhou Zhang
|
28a19e494b
|
Fix lint (#15281)
|
2025-12-16 13:36:01 -08:00 |
|
Baizhou Zhang
|
0261c4aff7
|
[misc] Upgrade cutedsl to 4.3.1 (#14857)
|
2025-12-16 12:11:56 -08:00 |
|
Baizhou Zhang
|
c843419562
|
Remove duplicate bs=1 in nightly benchmark (#15162)
|
2025-12-15 22:09:22 -08:00 |
|
Baizhou Zhang
|
ab3ffd1c8e
|
Add nightly accuracy test for DeepSeek V3.2 (#14935)
|
2025-12-13 12:11:16 -08:00 |
|
Baizhou Zhang
|
8698867479
|
[CI]Add gb200 runner back (#15024)
|
2025-12-12 20:19:34 -08:00 |
|
Baizhou Zhang
|
7dcad45cad
|
[CI] Temp disable gb200 test (#14865)
|
2025-12-10 19:05:16 -08:00 |
|
Baizhou Zhang
|
e5201bda34
|
[CI] Unblock gb200 cutedsl test (#14469)
|
2025-12-08 17:58:25 -08:00 |
|
Baizhou Zhang
|
6799847ebf
|
[CI]Unblock and split spec v2+dp test (#14551)
|
2025-12-07 17:39:25 -08:00 |
|
Baizhou Zhang
|
673c11ba73
|
[Minor] Temporarily skipping deepep large mtp test (#14586)
|
2025-12-07 13:59:16 -08:00 |
|
Baizhou Zhang
|
9dfa01a435
|
[Misc]Register and refactor some environs for dpsk-fp4 and DeepEp (#14538)
|
2025-12-06 12:29:16 -08:00 |
|
Baizhou Zhang
|
bc388471d2
|
[1/n] Fix hanging during DeepGemm Warmup (#14493)
|
2025-12-06 10:44:02 -08:00 |
|
Baizhou Zhang
|
42fcf5438f
|
Revert "tiny remove deprecated endpoint call" (#14533)
|
2025-12-05 23:48:54 -08:00 |
|
Baizhou Zhang
|
80a575e4e8
|
Add YAMY1234 to CI Permission (#14475)
|
2025-12-04 21:25:49 -08:00 |
|
 Baizhou ZhangandXinyuan Tong
|
7e78825d5a
|
[Tiny]Small fixes in deepseek v32 doc (#14372)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2025-12-03 11:35:40 -08:00 |
|
Baizhou Zhang
|
4bcc5879af
|
[Doc] Fix DeepSeek V32 Doc (#14336)
|
2025-12-02 21:06:55 -08:00 |
|
Baizhou Zhang
|
922054079c
|
[Doc] Update DeepSeek-V3.2 document (#14321)
|
2025-12-02 18:19:39 -08:00 |
|
Baizhou Zhang
|
03888b9de5
|
[Minor] Upgrade cutedsl version in Dockerfile (#13968)
|
2025-12-01 17:15:26 -08:00 |
|
Baizhou Zhang
|
eb5008846a
|
[CI] Fix test_deepep_large.py (#14247)
|
2025-12-01 15:18:48 -08:00 |
|
Baizhou Zhang
|
f1115cf58d
|
Revert "[Minor]Raise Error when deepep num dispatch token per rank is smaller than cuda graph bs" (#14171)
|
2025-11-30 12:49:46 -08:00 |
|
Baizhou Zhang
|
7b03cc6482
|
[Minor]Raise Error when deepep num dispatch token per rank is smaller than cuda graph bs (#14065)
|
2025-11-30 10:11:42 -08:00 |
|
Baizhou Zhang
|
051ad83347
|
[chore] Arrange NV packages in Dockerfile (#13749)
|
2025-11-27 18:08:27 -05:00 |
|
Baizhou Zhang
|
7ab548ef64
|
[2/2] Refactor DeepGeem requant for FP8 FusedMoE on Blackwell (#13960)
|
2025-11-27 09:00:26 -05:00 |
|
Baizhou Zhang
|
8a9b8b8457
|
Revert "Fix nightly test failures: NSA indexer dtype and CPP radix cache init" (#14015)
|
2025-11-26 10:45:23 -08:00 |
|
Baizhou Zhang
|
873382a910
|
[Tiny]Upgrade README for sgl-kernel (#13945)
|
2025-11-25 16:46:30 -08:00 |
|
Baizhou Zhang
|
808b6dfdea
|
[Minor] Fix lint (#13938)
|
2025-11-25 10:57:23 -08:00 |
|
Baizhou Zhang
|
04b52fa8d6
|
[chore]Upgrade flashinfer to 0.5.3 (#13751)
|
2025-11-23 23:38:36 -08:00 |
|
 Baizhou Zhangandfy1214
|
4683e244fe
|
[1/2] Refactor DeepGeem requant for FP8 Linear on Blackwell (#13601)
Co-authored-by: fy1214
|
2025-11-23 16:07:56 -08:00 |
|
Baizhou Zhang
|
c9bd1aca32
|
[CI] Tiny refactoring sgl-kernel tests (#13813)
|
2025-11-23 12:45:17 -08:00 |
|
Baizhou Zhang
|
8bfce9b08d
|
[Tiny] Renaming environ for NVFP4 dispatch (#13756)
|
2025-11-22 00:05:20 -08:00 |
|
Baizhou Zhang
|
9f59194f29
|
[Fix] Fix DeepSeek V3 MTP on B200 (#13548)
|
2025-11-18 16:30:26 -08:00 |
|
Baizhou Zhang
|
10969ae4be
|
[chore] Disable ccache for sgl-kernel release (#13541)
|
2025-11-18 14:28:58 -08:00 |
|
Baizhou Zhang
|
85ae508e8b
|
Add bfloat16 tuned fused moe config for Dpsk-MTP layer on B200 (#13455)
|
2025-11-17 17:44:31 -08:00 |
|
Baizhou Zhang
|
d64dd3e18e
|
[Tiny]Fix 1-gpu nightly test bugs (#13389)
|
2025-11-16 15:54:17 -08:00 |
|
Baizhou Zhang
|
3ccd7fa669
|
[CI] Fix B200 CI (#13387)
|
2025-11-16 15:13:57 -08:00 |
|
Baizhou Zhang
|
10285ec204
|
[Misc]Add date to cu13 dev image tag (#13316)
|
2025-11-14 19:42:50 -08:00 |
|
Baizhou Zhang
|
8ece99a9dc
|
[CI] Update job dependency and move dpsk v3.2 tests to 8-gpu suite (#12942)
|
2025-11-11 23:16:47 -08:00 |
|
Baizhou Zhang
|
99e25805f5
|
[Fix] Fix nan error for large scale ep (#12866)
|
2025-11-11 14:44:57 -08:00 |
|
Baizhou Zhang
|
5f02b918ec
|
[Fix] Fix trtllm-mla backend when chunked prefix cache is disabled (#12361)
|
2025-11-08 15:10:25 -08:00 |
|
Baizhou Zhang
|
e039ff382c
|
[CI] Fix huggingface access for test_flash_attention_4.py (#12846)
|
2025-11-07 20:07:06 -08:00 |
|
Baizhou Zhang
|
bb6a21cd99
|
[Fix]Tiny fix in Dockerfile (#12748)
|
2025-11-05 21:04:33 -08:00 |
|
Baizhou Zhang
|
3c0a6df82d
|
[chore] Fix triton installation for cu13 image (#12742)
|
2025-11-05 20:20:18 -08:00 |
|
Baizhou Zhang
|
9a954982de
|
[chore] SGLang tag management in Dockerfile (#12734)
|
2025-11-05 19:03:26 -08:00 |
|
Baizhou Zhang
|
9ec6031d70
|
[chore]Remove dockerfile from target file of bump kernel version (#12728)
|
2025-11-05 18:01:19 -08:00 |
|
Baizhou Zhang
|
7c45b8b4bb
|
[CI] Fix qwen3-vl lora nightly ci (#12708)
|
2025-11-05 11:00:13 -08:00 |
|
Baizhou Zhang
|
d22d044734
|
Revert "Enable memory saver for hybrid model" (#12648)
|
2025-11-04 16:22:06 -08:00 |
|
Baizhou Zhang
|
42889acbd0
|
[hotfix] Fix deepep w4a8 bug (#12642)
|
2025-11-04 13:55:59 -08:00 |
|
Baizhou Zhang
|
15efbcb4e7
|
[chore] Fix update_kernel_whl_index script for multiple cuda version (#12519)
|
2025-11-03 16:34:14 -08:00 |
|
Baizhou Zhang
|
6e29446e45
|
[hotfix] Remove flashinfer-jit-cache from pyproject (#12530)
|
2025-11-02 22:11:05 -08:00 |
|
Baizhou Zhang
|
9a512cf95b
|
[CI] Move some Lora/Deterministic CI tests to nightly (#12507)
|
2025-11-01 19:54:22 -07:00 |
|
Baizhou Zhang
|
2b7bf11bd2
|
[Hotfix] Remove extra comment in sgl-kernel README (#12500)
|
2025-11-01 12:22:45 -07:00 |
|
Baizhou Zhang
|
566ade0388
|
[CI] Build aarch64 kernels for sgl-kernel test (#12480)
|
2025-11-01 11:55:42 -07:00 |
|
Baizhou Zhang
|
5f98b7fe61
|
[CI] Fix kernel installation on aarch runners (#12475)
|
2025-10-31 14:25:27 -07:00 |
|