Commit Graph

9148 Commits

Author SHA1 Message Date
Yi Zhong
08fcda2f63 add the fa4 mm backend and varlen func (#13539)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-01-23 23:12:06 +08:00
akhilg-nv
2fb328109f [DeepSeek V3.2] Enable trtllm NSA with bf16 kvcache (#16758)
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
2026-01-23 20:26:21 +08:00
Nicolas Castet
48e9daadff Support symmetric memory pre-allocation to avoid fragmentation (#17089) 2026-01-23 17:57:04 +08:00
Alison Shao
b23470e95a Fix CI install failure when rerunning tests via workflow_dispatch (#17612) 2026-01-23 00:04:16 -08:00
siyu
17349168bb set cooldown_interval_minutes to 0 for liusy58 (#17637) 2026-01-22 23:36:05 -08:00
Yuzhen Zhou
2169025b77 turn off dit_layerwise_offload for wan on rocm (#17569) 2026-01-23 15:22:42 +08:00
Hexq0210
69470dbc1f [NPU] update doc for Ascend NPU (#17621) 2026-01-23 15:21:00 +08:00
Even Zhou
69ac8b58f7 [NPU] [CI] temporarily disable mtp test (#17614) 2026-01-23 15:17:31 +08:00
Minglei Zhu
2b2f317383 fix gpt-oss launch failure with piecewise cuda graph (#17532) 2026-01-22 22:41:38 -08:00
Alison Shao
d7dd0b8832 Re-enable unit-test-deepep-8-gpu and unit-test-backend-4-gpu-gb200 (#17438) 2026-01-23 14:31:44 +08:00
Lianmin Zheng
56e6652d1d Lazy import torchao (#17626) 2026-01-22 22:04:51 -08:00
Bingxu Chen
50a2e4345a [AMD CI] Add 2-GPU sgl-kernel Tests (#17555)
Co-authored-by: YC Tseng <yctseng@amd.com>
2026-01-22 21:48:52 -08:00
JiaruiChang5268
c0b5a180fe [NPU]bugfix: fix for dsv3.2 and dsvl2 (#17007)
Co-authored-by: Hexq0210 <893781835@qq.com>
Co-authored-by: liupeng374 <782420244@qq.com>
Co-authored-by: cy <chenyang08056032@163.com>
2026-01-23 11:15:15 +08:00
Ke Bao
7ace64d1d8 Update mamba env setting (#17566) 2026-01-23 11:02:32 +08:00
siyu
62e6a749b0 Skip mm feature pool init to avoid EPD OOM (#16388) 2026-01-23 10:53:45 +08:00
YC Tseng
04a10c9bc2 [AMD] CI - migrate perf test and fix stage-b-test-1-gpu-amd (#17340)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
Co-authored-by: michaelzhang-ai <michaelzhang.ai@users.noreply.github.com>
2026-01-22 18:45:05 -08:00
amote-i
0fec8820d1 update dependence docs of npu (#17573) 2026-01-23 09:13:36 +08:00
Alison Shao
1e8e0cca2c Update test README with CI registry documentation and 5090/H100 guidance (#17368) 2026-01-22 16:29:08 -08:00
MMuzzammil1
2399af5557 Bugfix: Writing to storage when write-back method is chosen (#14718) 2026-01-22 15:08:25 -08:00
hxie
13f88045b3 configuration file support and nixl integration augmentation for hicache-storage-backend-extra-config (#16602) 2026-01-22 14:31:48 -08:00
Jacob Gordon
a296c99ff4 refactor(benchmark): prevents variable shadowing (#17607) 2026-01-22 17:00:11 -05:00
wufann
a921029b97 [AMD] Support ds3.2 on gfx942 platform (#17504)
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
2026-01-22 13:57:08 -08:00
Jacob Gordon
15b511771d refactor(codespell): corrects typos covered up by whitelist (#17601) 2026-01-22 11:45:47 -08:00
Jacob Gordon
2c86a065db types: preps SessionReqNode for IDE-driven renames (#17602) 2026-01-22 11:36:23 -08:00
Baizhou Zhang
283a2daeaa [hotfix] Reenable all reduce fusion on sm100 (#17591) 2026-01-22 23:36:38 +08:00
Yuhao Yang
f7a0bcda1e model: step3-vl-10b (#17513) 2026-01-22 23:15:08 +08:00
Jincong Chen
72f790bf6f [BUGFIX] Skip Mamba Cache Slot 0 to Avoid Using Dummy Cache (#17404) 2026-01-22 23:08:38 +08:00
Jacob Gordon
42523d0364 ci(pre-commit): avoids extraneous codespell exclusions (#17590) 2026-01-22 09:55:23 -05:00
Xiaoyu Zhang
5324027007 [Diffusion] Make the apply_qknorm function easier to use (#17537) 2026-01-22 22:32:15 +08:00
triple-mu
3705f90629 [diffusion] model: optimize torch.compile (#17472) 2026-01-22 22:05:47 +08:00
chenxu214
5d299c25c0 [NPU] bugfix with Kimi-k2 and bge-reranker-v2 model (#17478)
Co-authored-by: amote-i <49533125+amote-i@users.noreply.github.com>
Co-authored-by: cy <chenyang08056032@163.com>
2026-01-22 22:02:05 +08:00
chenxu214
a4dc432587 Change naming for graph mode on multiplatform (#17469) 2026-01-22 19:25:06 +08:00
Minglei Zhu
419bbcee10 refactor Qwen3-Next with a new RadixLinearAttention (#17373) 2026-01-22 17:42:06 +08:00
zhangheng
f33022d039 [RadixTree][3/N Refactor]:Support unified insert/evict params (#17401) 2026-01-22 17:36:31 +08:00
Shangming Cai
2262c5c9b5 Add ZhengWG to CI_Permission (#17572) 2026-01-22 17:24:06 +08:00
Baizhou Zhang
8dae6ec03c Add xyjixyjixyji to CI_Permission (#17559) 2026-01-21 23:27:25 -08:00
Chandrakant Khandelwal
61abff66c1 [NPU] [Bug Fix] Fix typo in npu device check in gpt_oss.py (#17553) 2026-01-21 23:26:35 -08:00
Zaili Wang
6a8f68b6d2 [Fix] fix device orientation for image processor (#15859) 2026-01-22 15:03:10 +08:00
Zaili Wang
672eb37534 [CPU][Fix CI] Solidate torch version for sgl-kernel-cpu and fix device orientation error (#17460) 2026-01-22 14:04:50 +08:00
Baizhou Zhang
e2d33531f3 [Kernel] Little refactor of flashinfer allreduce norm fusion (#17474) 2026-01-22 13:31:57 +08:00
Chi McIsaac
71482dd171 [diffusion] feat: enable passing Cache‑DiT config for diffusers backend (#16662)
Signed-off-by: Chi <chixie.mcisaac@gmail.com>
Signed-off-by: qimcis <chixie.mcisaac@gmail.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-01-22 13:13:34 +08:00
YC Tseng
17807caf82 [AMD] fix amd ci dpskv32 (#17432)
Co-authored-by: michaelzhang-ai <michaelzhang.ai@users.noreply.github.com>
2026-01-21 20:34:24 -08:00
Baizhou Zhang
fafa171529 [hotfix] Fixes on cuda 13 docker image (#17541)
Co-authored-by: iforgetmyname <iforgetmyname@users.noreply.github>
2026-01-22 12:29:55 +08:00
Yingchun Lai
a5bbcda968 fix: prefer to use max_completion_tokens rather than max_tokens (#17516) 2026-01-21 20:22:20 -08:00
Serge Panev
e95668abc7 [NVIDIA] Fix CUDA arch requirement in nvfp4 cast (#12581)
Signed-off-by: Serge Panev <spanev@nvidia.com>
Co-authored-by: Fan Yin <1106310035@qq.com>
2026-01-21 20:21:11 -08:00
Baizhou Zhang
3373545b9f [HotFix]Fix dtype mismatch in nsa indexer on AMD device (#17518) 2026-01-22 12:17:02 +08:00
Piotr Mazurek
d6e2b88288 Add Liquid Foundation Model (LFM2) (#16890) 2026-01-22 11:11:20 +08:00
Kangyan-Zhou
f2ae066a6b Update release-branch-cut.yml for actions: write (#17539) 2026-01-21 18:03:27 -08:00
cen121212
0c2993eed0 Optimize Qwen3-VL video memory usage (#16366) 2026-01-22 09:10:08 +08:00
Xiaoyu Zhang
590969ee9c [Diffusion] Support select fa2 backend in hopper (#17514) 2026-01-22 08:23:53 +08:00