Commit Graph

9103 Commits

Author SHA1 Message Date
Baizhou Zhang
3373545b9f [HotFix]Fix dtype mismatch in nsa indexer on AMD device (#17518) 2026-01-22 12:17:02 +08:00
Piotr Mazurek
d6e2b88288 Add Liquid Foundation Model (LFM2) (#16890) 2026-01-22 11:11:20 +08:00
Kangyan-Zhou
f2ae066a6b Update release-branch-cut.yml for actions: write (#17539) 2026-01-21 18:03:27 -08:00
cen121212
0c2993eed0 Optimize Qwen3-VL video memory usage (#16366) 2026-01-22 09:10:08 +08:00
Xiaoyu Zhang
590969ee9c [Diffusion] Support select fa2 backend in hopper (#17514) 2026-01-22 08:23:53 +08:00
Alison Shao
e6ccb2949b Increase wait-for-stage timeouts to handle long queue times (#17536) 2026-01-21 16:10:52 -08:00
Alison Shao
d6ea2c529c Fix import path for UnquantizedLinearMethod in test (#17529) 2026-01-21 15:34:02 -08:00
Alison Shao
9be2a3a9a3 Remove test_gpt_oss_4gpu.py from __not_in_ci__ (keep in per-commit-4-gpu) (#17534) 2026-01-21 15:31:03 -08:00
Lianmin Zheng
b74a57a8d9 [Auto Sync] Update detokenizer_manager.py, io_struct.py, mu... (20260120) (#17442)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Wangfan Fu <wangfan@x.ai>
2026-01-21 14:48:32 -08:00
DarkSharpness
95f59c13fd [Chore] include all jit files in building packages (#17493) 2026-01-21 14:48:02 -08:00
Alison Shao
85d9af51da Temporarily disable flaky test_gpt_oss_4gpu.py on B200 (#17528) 2026-01-21 14:01:04 -08:00
Jacob Gordon
858f317f13 ci(codespell): centralizes list of ignorable words (#17524) 2026-01-21 12:29:14 -08:00
Lingjun Wen
cf89351691 [new-model] Add support for Cohere2ForCausalLM behind Command-A and Command-R Models (#16927) 2026-01-21 12:28:33 -08:00
Lianmin Zheng
1fdf5cac39 [Auto Sync] Update environ.py, fp8.py (20260121) (#17486)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Binyao Jiang <byjiang1996@gmail.com>
2026-01-21 12:04:09 -08:00
Jacob Gordon
cda43ffa4d ci: avoids duplication of codespell config (#17519) 2026-01-21 12:02:37 -08:00
Yunmeng
390898545e [Misc] Fix argument help string formatting (#17416) 2026-01-21 09:25:32 -08:00
YC Tseng
b827e9d381 [AMD] CI - Fix sgl-kernel unittest (#17490) 2026-01-21 09:23:28 -08:00
Ke Bao
d725487dc8 Disable swa memory for gpt-oss with spec (#17517) 2026-01-22 01:04:19 +08:00
JiaruiChang5268
a95c9f5b81 [NPU] Remove paged attention & Change fia to default attention (#17394)
Co-authored-by: Liwansi <62291011+Liwansi@users.noreply.github.com>
Co-authored-by: chenxu214 <justin_cc2025@163.com>
Co-authored-by: chenyang08056032 <chenyang08056032@163.com>
2026-01-21 23:58:25 +08:00
Xiaoyu Zhang
19089aa431 [Diffusion] Refactor diffusion is_cuda check (#17498) 2026-01-21 23:02:24 +08:00
Qiaolin Yu
4f6f5d25c8 Support fa4 decoding (#16034) 2026-01-21 22:54:02 +08:00
Yi Zhong
458fe5a337 [docs] Show user the fastAPI docs available (#17510)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
2026-01-21 14:26:25 +00:00
b8zhong
2ff0880a0e [Fix] GLM 4.7 + NVFP4 + MTP (#17166) 2026-01-21 21:34:18 +08:00
Zhu Yuhua
2c1b164a92 [diffusion] improve: skip negative prompt encoding when guidance_scale <= 1.0 or negative_prompt is None (#16919)
Signed-off-by: zhuyuhua-v <yuhzhu@amd.com>
2026-01-21 20:01:55 +08:00
strgrb
bcc6d84f93 Use fused_sigmoid_gating_delta_rule_update_kernel for KDA (#17108) 2026-01-21 19:24:29 +08:00
Ke Bao
a618202fc7 Tiny refine swa kv cache free (#17417) 2026-01-21 19:20:05 +08:00
siyu
7520b92927 Support EPD error handling (#16670)
Co-authored-by: ZhengWG <zwg0606@gmail.com>
2026-01-21 18:47:00 +08:00
Fan Lin
e7224e9681 [diffusion] fix: fix the LoRA weights mismatch caused by weights packing (#17355) 2026-01-21 18:17:16 +08:00
HuangJi
e776239afd [diffusion] feat: support SageSparseLinearAttention attention backend (#17399) 2026-01-21 18:13:51 +08:00
Yi Zhang
1b97fa769b [BUGFIX] fix value oom in radix tree (#17400) 2026-01-21 17:12:57 +08:00
Yi Zhang
236772c0e1 [RadixTree][2/N Refactor]: swa cache init tiny refactor (#17397) 2026-01-21 15:48:30 +08:00
Sam Shleifer
0d49b13fdd Fix circular import in quantization modules (#17372) 2026-01-21 15:47:09 +08:00
blahblah
0a7a2017a0 [diffusion] refactor: refactor and simplify teacache for cachabledit and wanvideo (#16396)
Co-authored-by: Brain97 <Brain97@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: blahblah <blahblah>
2026-01-21 15:42:45 +08:00
amote-i
0a9099e137 update ascend docs (#17457) 2026-01-21 15:36:26 +08:00
Alison Shao
0050c476fd Add job-level timeout for weekly test workflow (#17462) 2026-01-20 22:45:39 -08:00
Baizhou Zhang
8251a74d5f [Tiny] Backward compatibility for fp4 gemm flags (#17466) 2026-01-21 14:34:40 +08:00
Baizhou Zhang
a54d75bf2e [Fix] Set fa3 as default MHA backend on Hopper (#17425) 2026-01-21 13:54:09 +08:00
Alison Shao
3321eb4efa Fix pr-test-finish to fail when wait-for-stage jobs fail (#17465) 2026-01-20 20:50:27 -08:00
Kangyan-Zhou
be5121b452 Fix NSA indexer in the nightly test (#17452) 2026-01-20 19:58:13 -08:00
Baizhou Zhang
c3f9c30f99 [Minor] Change lora_target_modules to "all" in CI tests (#17386) 2026-01-21 11:46:36 +08:00
Fan Lin
54a821794e [diffusion] fix: fix the bug of output_path not taking effect when generate (#17293) 2026-01-21 11:39:51 +08:00
khalilzhk
aca354bcb3 [NPU] remove features supported on Ascend NPU (#17455) 2026-01-21 11:00:04 +08:00
Yinghai Lu
aea57b33c6 [scheduler] Clear MM data of finished batch (#17251)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2026-01-20 17:45:38 -08:00
Alison Shao
823a046e8f Add hybrid parallelism test to nightly CI (#17444) 2026-01-20 17:43:50 -08:00
Alison Shao
648aab0ce3 Fix wait-for-stage jobs running when call-gate fails (#17443) 2026-01-20 17:40:05 -08:00
ishandhanani
1e309030e3 update urllib3 and gpgv Dockerfile (#17439) 2026-01-20 14:47:20 -08:00
Binyao Jiang
6092721594 [Piecewise] Fix PCG issue for multimodal and embedding model that wraps language_model (#17290) 2026-01-20 14:06:06 -08:00
Binyao Jiang
38c233fd04 [Piecewise] Support PCG weak_ref_tensor cuda kernel on AMD (#17291) 2026-01-20 14:05:32 -08:00
Lianmin Zheng
20ed3822bb [Auto Sync] Update piecewise_cuda_graph_runner.py (20260119) (#17313)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Binyao Jiang <byjiang1996@gmail.com>
2026-01-20 14:05:05 -08:00
zijiexia
4ecd9afde9 [Docs] Rename SGLang Router to SGLang Model Gateway (#17436) 2026-01-20 12:31:10 -08:00