Yi Zhong
|
08fcda2f63
|
add the fa4 mm backend and varlen func (#13539)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-01-23 23:12:06 +08:00 |
|
akhilg-nv
|
2fb328109f
|
[DeepSeek V3.2] Enable trtllm NSA with bf16 kvcache (#16758)
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
|
2026-01-23 20:26:21 +08:00 |
|
Nicolas Castet
|
48e9daadff
|
Support symmetric memory pre-allocation to avoid fragmentation (#17089)
|
2026-01-23 17:57:04 +08:00 |
|
Alison Shao
|
b23470e95a
|
Fix CI install failure when rerunning tests via workflow_dispatch (#17612)
|
2026-01-23 00:04:16 -08:00 |
|
siyu
|
17349168bb
|
set cooldown_interval_minutes to 0 for liusy58 (#17637)
|
2026-01-22 23:36:05 -08:00 |
|
Yuzhen Zhou
|
2169025b77
|
turn off dit_layerwise_offload for wan on rocm (#17569)
|
2026-01-23 15:22:42 +08:00 |
|
Hexq0210
|
69470dbc1f
|
[NPU] update doc for Ascend NPU (#17621)
|
2026-01-23 15:21:00 +08:00 |
|
Even Zhou
|
69ac8b58f7
|
[NPU] [CI] temporarily disable mtp test (#17614)
|
2026-01-23 15:17:31 +08:00 |
|
Minglei Zhu
|
2b2f317383
|
fix gpt-oss launch failure with piecewise cuda graph (#17532)
|
2026-01-22 22:41:38 -08:00 |
|
Alison Shao
|
d7dd0b8832
|
Re-enable unit-test-deepep-8-gpu and unit-test-backend-4-gpu-gb200 (#17438)
|
2026-01-23 14:31:44 +08:00 |
|
Lianmin Zheng
|
56e6652d1d
|
Lazy import torchao (#17626)
|
2026-01-22 22:04:51 -08:00 |
|
Bingxu Chen
|
50a2e4345a
|
[AMD CI] Add 2-GPU sgl-kernel Tests (#17555)
Co-authored-by: YC Tseng <yctseng@amd.com>
|
2026-01-22 21:48:52 -08:00 |
|
JiaruiChang5268
|
c0b5a180fe
|
[NPU]bugfix: fix for dsv3.2 and dsvl2 (#17007)
Co-authored-by: Hexq0210 <893781835@qq.com>
Co-authored-by: liupeng374 <782420244@qq.com>
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-01-23 11:15:15 +08:00 |
|
Ke Bao
|
7ace64d1d8
|
Update mamba env setting (#17566)
|
2026-01-23 11:02:32 +08:00 |
|
siyu
|
62e6a749b0
|
Skip mm feature pool init to avoid EPD OOM (#16388)
|
2026-01-23 10:53:45 +08:00 |
|
YC Tseng
|
04a10c9bc2
|
[AMD] CI - migrate perf test and fix stage-b-test-1-gpu-amd (#17340)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
Co-authored-by: michaelzhang-ai <michaelzhang.ai@users.noreply.github.com>
|
2026-01-22 18:45:05 -08:00 |
|
amote-i
|
0fec8820d1
|
update dependence docs of npu (#17573)
|
2026-01-23 09:13:36 +08:00 |
|
Alison Shao
|
1e8e0cca2c
|
Update test README with CI registry documentation and 5090/H100 guidance (#17368)
|
2026-01-22 16:29:08 -08:00 |
|
MMuzzammil1
|
2399af5557
|
Bugfix: Writing to storage when write-back method is chosen (#14718)
|
2026-01-22 15:08:25 -08:00 |
|
hxie
|
13f88045b3
|
configuration file support and nixl integration augmentation for hicache-storage-backend-extra-config (#16602)
|
2026-01-22 14:31:48 -08:00 |
|
Jacob Gordon
|
a296c99ff4
|
refactor(benchmark): prevents variable shadowing (#17607)
|
2026-01-22 17:00:11 -05:00 |
|
wufann
|
a921029b97
|
[AMD] Support ds3.2 on gfx942 platform (#17504)
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
|
2026-01-22 13:57:08 -08:00 |
|
Jacob Gordon
|
15b511771d
|
refactor(codespell): corrects typos covered up by whitelist (#17601)
|
2026-01-22 11:45:47 -08:00 |
|
Jacob Gordon
|
2c86a065db
|
types: preps SessionReqNode for IDE-driven renames (#17602)
|
2026-01-22 11:36:23 -08:00 |
|
Baizhou Zhang
|
283a2daeaa
|
[hotfix] Reenable all reduce fusion on sm100 (#17591)
|
2026-01-22 23:36:38 +08:00 |
|
Yuhao Yang
|
f7a0bcda1e
|
model: step3-vl-10b (#17513)
|
2026-01-22 23:15:08 +08:00 |
|
Jincong Chen
|
72f790bf6f
|
[BUGFIX] Skip Mamba Cache Slot 0 to Avoid Using Dummy Cache (#17404)
|
2026-01-22 23:08:38 +08:00 |
|
Jacob Gordon
|
42523d0364
|
ci(pre-commit): avoids extraneous codespell exclusions (#17590)
|
2026-01-22 09:55:23 -05:00 |
|
Xiaoyu Zhang
|
5324027007
|
[Diffusion] Make the apply_qknorm function easier to use (#17537)
|
2026-01-22 22:32:15 +08:00 |
|
triple-mu
|
3705f90629
|
[diffusion] model: optimize torch.compile (#17472)
|
2026-01-22 22:05:47 +08:00 |
|
chenxu214
|
5d299c25c0
|
[NPU] bugfix with Kimi-k2 and bge-reranker-v2 model (#17478)
Co-authored-by: amote-i <49533125+amote-i@users.noreply.github.com>
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-01-22 22:02:05 +08:00 |
|
chenxu214
|
a4dc432587
|
Change naming for graph mode on multiplatform (#17469)
|
2026-01-22 19:25:06 +08:00 |
|
Minglei Zhu
|
419bbcee10
|
refactor Qwen3-Next with a new RadixLinearAttention (#17373)
|
2026-01-22 17:42:06 +08:00 |
|
zhangheng
|
f33022d039
|
[RadixTree][3/N Refactor]:Support unified insert/evict params (#17401)
|
2026-01-22 17:36:31 +08:00 |
|
Shangming Cai
|
2262c5c9b5
|
Add ZhengWG to CI_Permission (#17572)
|
2026-01-22 17:24:06 +08:00 |
|
Baizhou Zhang
|
8dae6ec03c
|
Add xyjixyjixyji to CI_Permission (#17559)
|
2026-01-21 23:27:25 -08:00 |
|
Chandrakant Khandelwal
|
61abff66c1
|
[NPU] [Bug Fix] Fix typo in npu device check in gpt_oss.py (#17553)
|
2026-01-21 23:26:35 -08:00 |
|
Zaili Wang
|
6a8f68b6d2
|
[Fix] fix device orientation for image processor (#15859)
|
2026-01-22 15:03:10 +08:00 |
|
Zaili Wang
|
672eb37534
|
[CPU][Fix CI] Solidate torch version for sgl-kernel-cpu and fix device orientation error (#17460)
|
2026-01-22 14:04:50 +08:00 |
|
Baizhou Zhang
|
e2d33531f3
|
[Kernel] Little refactor of flashinfer allreduce norm fusion (#17474)
|
2026-01-22 13:31:57 +08:00 |
|
Chi McIsaac
|
71482dd171
|
[diffusion] feat: enable passing Cache‑DiT config for diffusers backend (#16662)
Signed-off-by: Chi <chixie.mcisaac@gmail.com>
Signed-off-by: qimcis <chixie.mcisaac@gmail.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-01-22 13:13:34 +08:00 |
|
YC Tseng
|
17807caf82
|
[AMD] fix amd ci dpskv32 (#17432)
Co-authored-by: michaelzhang-ai <michaelzhang.ai@users.noreply.github.com>
|
2026-01-21 20:34:24 -08:00 |
|
Baizhou Zhang
|
fafa171529
|
[hotfix] Fixes on cuda 13 docker image (#17541)
Co-authored-by: iforgetmyname <iforgetmyname@users.noreply.github>
|
2026-01-22 12:29:55 +08:00 |
|
Yingchun Lai
|
a5bbcda968
|
fix: prefer to use max_completion_tokens rather than max_tokens (#17516)
|
2026-01-21 20:22:20 -08:00 |
|
Serge Panev
|
e95668abc7
|
[NVIDIA] Fix CUDA arch requirement in nvfp4 cast (#12581)
Signed-off-by: Serge Panev <spanev@nvidia.com>
Co-authored-by: Fan Yin <1106310035@qq.com>
|
2026-01-21 20:21:11 -08:00 |
|
Baizhou Zhang
|
3373545b9f
|
[HotFix]Fix dtype mismatch in nsa indexer on AMD device (#17518)
|
2026-01-22 12:17:02 +08:00 |
|
Piotr Mazurek
|
d6e2b88288
|
Add Liquid Foundation Model (LFM2) (#16890)
|
2026-01-22 11:11:20 +08:00 |
|
Kangyan-Zhou
|
f2ae066a6b
|
Update release-branch-cut.yml for actions: write (#17539)
|
2026-01-21 18:03:27 -08:00 |
|
cen121212
|
0c2993eed0
|
Optimize Qwen3-VL video memory usage (#16366)
|
2026-01-22 09:10:08 +08:00 |
|
Xiaoyu Zhang
|
590969ee9c
|
[Diffusion] Support select fa2 backend in hopper (#17514)
|
2026-01-22 08:23:53 +08:00 |
|