Commit Graph

11116 Commits

Author SHA1 Message Date
Minglei Zhu
419bbcee10 refactor Qwen3-Next with a new RadixLinearAttention (#17373) 2026-01-22 17:42:06 +08:00
zhangheng
f33022d039 [RadixTree][3/N Refactor]:Support unified insert/evict params (#17401) 2026-01-22 17:36:31 +08:00
Shangming Cai
2262c5c9b5 Add ZhengWG to CI_Permission (#17572) 2026-01-22 17:24:06 +08:00
Baizhou Zhang
8dae6ec03c Add xyjixyjixyji to CI_Permission (#17559) 2026-01-21 23:27:25 -08:00
Chandrakant Khandelwal
61abff66c1 [NPU] [Bug Fix] Fix typo in npu device check in gpt_oss.py (#17553) 2026-01-21 23:26:35 -08:00
Zaili Wang
6a8f68b6d2 [Fix] fix device orientation for image processor (#15859) 2026-01-22 15:03:10 +08:00
Zaili Wang
672eb37534 [CPU][Fix CI] Solidate torch version for sgl-kernel-cpu and fix device orientation error (#17460) 2026-01-22 14:04:50 +08:00
Baizhou Zhang
e2d33531f3 [Kernel] Little refactor of flashinfer allreduce norm fusion (#17474) 2026-01-22 13:31:57 +08:00
Chi McIsaac
71482dd171 [diffusion] feat: enable passing Cache‑DiT config for diffusers backend (#16662)
Signed-off-by: Chi <chixie.mcisaac@gmail.com>
Signed-off-by: qimcis <chixie.mcisaac@gmail.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-01-22 13:13:34 +08:00
YC Tseng
17807caf82 [AMD] fix amd ci dpskv32 (#17432)
Co-authored-by: michaelzhang-ai <michaelzhang.ai@users.noreply.github.com>
2026-01-21 20:34:24 -08:00
Baizhou Zhang
fafa171529 [hotfix] Fixes on cuda 13 docker image (#17541)
Co-authored-by: iforgetmyname <iforgetmyname@users.noreply.github>
2026-01-22 12:29:55 +08:00
Yingchun Lai
a5bbcda968 fix: prefer to use max_completion_tokens rather than max_tokens (#17516) 2026-01-21 20:22:20 -08:00
Serge Panev
e95668abc7 [NVIDIA] Fix CUDA arch requirement in nvfp4 cast (#12581)
Signed-off-by: Serge Panev <spanev@nvidia.com>
Co-authored-by: Fan Yin <1106310035@qq.com>
2026-01-21 20:21:11 -08:00
Baizhou Zhang
3373545b9f [HotFix]Fix dtype mismatch in nsa indexer on AMD device (#17518) 2026-01-22 12:17:02 +08:00
Piotr Mazurek
d6e2b88288 Add Liquid Foundation Model (LFM2) (#16890) 2026-01-22 11:11:20 +08:00
Kangyan-Zhou
f2ae066a6b Update release-branch-cut.yml for actions: write (#17539) 2026-01-21 18:03:27 -08:00
cen121212
0c2993eed0 Optimize Qwen3-VL video memory usage (#16366) 2026-01-22 09:10:08 +08:00
Xiaoyu Zhang
590969ee9c [Diffusion] Support select fa2 backend in hopper (#17514) 2026-01-22 08:23:53 +08:00
Alison Shao
e6ccb2949b Increase wait-for-stage timeouts to handle long queue times (#17536) 2026-01-21 16:10:52 -08:00
Alison Shao
d6ea2c529c Fix import path for UnquantizedLinearMethod in test (#17529) 2026-01-21 15:34:02 -08:00
Alison Shao
9be2a3a9a3 Remove test_gpt_oss_4gpu.py from __not_in_ci__ (keep in per-commit-4-gpu) (#17534) 2026-01-21 15:31:03 -08:00
Lianmin Zheng
b74a57a8d9 [Auto Sync] Update detokenizer_manager.py, io_struct.py, mu... (20260120) (#17442)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Wangfan Fu <wangfan@x.ai>
2026-01-21 14:48:32 -08:00
DarkSharpness
95f59c13fd [Chore] include all jit files in building packages (#17493) 2026-01-21 14:48:02 -08:00
Alison Shao
85d9af51da Temporarily disable flaky test_gpt_oss_4gpu.py on B200 (#17528) 2026-01-21 14:01:04 -08:00
Jacob Gordon
858f317f13 ci(codespell): centralizes list of ignorable words (#17524) 2026-01-21 12:29:14 -08:00
Lingjun Wen
cf89351691 [new-model] Add support for Cohere2ForCausalLM behind Command-A and Command-R Models (#16927) 2026-01-21 12:28:33 -08:00
Lianmin Zheng
1fdf5cac39 [Auto Sync] Update environ.py, fp8.py (20260121) (#17486)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Binyao Jiang <byjiang1996@gmail.com>
2026-01-21 12:04:09 -08:00
Jacob Gordon
cda43ffa4d ci: avoids duplication of codespell config (#17519) 2026-01-21 12:02:37 -08:00
Yunmeng
390898545e [Misc] Fix argument help string formatting (#17416) 2026-01-21 09:25:32 -08:00
YC Tseng
b827e9d381 [AMD] CI - Fix sgl-kernel unittest (#17490) 2026-01-21 09:23:28 -08:00
Ke Bao
d725487dc8 Disable swa memory for gpt-oss with spec (#17517) 2026-01-22 01:04:19 +08:00
JiaruiChang5268
a95c9f5b81 [NPU] Remove paged attention & Change fia to default attention (#17394)
Co-authored-by: Liwansi <62291011+Liwansi@users.noreply.github.com>
Co-authored-by: chenxu214 <justin_cc2025@163.com>
Co-authored-by: chenyang08056032 <chenyang08056032@163.com>
2026-01-21 23:58:25 +08:00
Xiaoyu Zhang
19089aa431 [Diffusion] Refactor diffusion is_cuda check (#17498) 2026-01-21 23:02:24 +08:00
Qiaolin Yu
4f6f5d25c8 Support fa4 decoding (#16034) 2026-01-21 22:54:02 +08:00
Yi Zhong
458fe5a337 [docs] Show user the fastAPI docs available (#17510)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
2026-01-21 14:26:25 +00:00
b8zhong
2ff0880a0e [Fix] GLM 4.7 + NVFP4 + MTP (#17166) 2026-01-21 21:34:18 +08:00
Zhu Yuhua
2c1b164a92 [diffusion] improve: skip negative prompt encoding when guidance_scale <= 1.0 or negative_prompt is None (#16919)
Signed-off-by: zhuyuhua-v <yuhzhu@amd.com>
2026-01-21 20:01:55 +08:00
strgrb
bcc6d84f93 Use fused_sigmoid_gating_delta_rule_update_kernel for KDA (#17108) 2026-01-21 19:24:29 +08:00
Ke Bao
a618202fc7 Tiny refine swa kv cache free (#17417) 2026-01-21 19:20:05 +08:00
siyu
7520b92927 Support EPD error handling (#16670)
Co-authored-by: ZhengWG <zwg0606@gmail.com>
2026-01-21 18:47:00 +08:00
Fan Lin
e7224e9681 [diffusion] fix: fix the LoRA weights mismatch caused by weights packing (#17355) 2026-01-21 18:17:16 +08:00
HuangJi
e776239afd [diffusion] feat: support SageSparseLinearAttention attention backend (#17399) 2026-01-21 18:13:51 +08:00
Yi Zhang
1b97fa769b [BUGFIX] fix value oom in radix tree (#17400) 2026-01-21 17:12:57 +08:00
Yi Zhang
236772c0e1 [RadixTree][2/N Refactor]: swa cache init tiny refactor (#17397) 2026-01-21 15:48:30 +08:00
Sam Shleifer
0d49b13fdd Fix circular import in quantization modules (#17372) 2026-01-21 15:47:09 +08:00
blahblah
0a7a2017a0 [diffusion] refactor: refactor and simplify teacache for cachabledit and wanvideo (#16396)
Co-authored-by: Brain97 <Brain97@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: blahblah <blahblah>
2026-01-21 15:42:45 +08:00
amote-i
0a9099e137 update ascend docs (#17457) 2026-01-21 15:36:26 +08:00
Alison Shao
0050c476fd Add job-level timeout for weekly test workflow (#17462) 2026-01-20 22:45:39 -08:00
Baizhou Zhang
8251a74d5f [Tiny] Backward compatibility for fp4 gemm flags (#17466) 2026-01-21 14:34:40 +08:00
Baizhou Zhang
a54d75bf2e [Fix] Set fa3 as default MHA backend on Hopper (#17425) 2026-01-21 13:54:09 +08:00