Commit Graph

9627 Commits

Author SHA1 Message Date
Thomas Wang
e20e6c28b9 [AMD] Fix accuracy issue when running TP4 dsv3 model with mtp (#18607)
Co-authored-by: YC Tseng <yctseng@amd.com>
Co-authored-by: kkHuang-amd <wunhuang@amd.com>
2026-02-12 01:13:16 -08:00
YC Tseng
d6f0ef677b [AMD] reset AMD image release time and reduce CI queue time (#18707) 2026-02-12 01:05:53 -08:00
Alison Shao
f20b1703ce [CI] Fix torchaudio/torchvision CUDA version mismatch (#18211) 2026-02-11 23:47:32 -08:00
chenxu214
1edc69be08 [Ascend]Support qwen3.5 (#18544)
This PR affects only the NPU. If any issues arise, please contact iforgetmyname.
2026-02-12 15:22:47 +08:00
Alan Kao
0305d12df2 [AMD] Enable release image build for ROCm 7.2.0 (#18698) 2026-02-11 23:16:54 -08:00
Vinh H. Pham
feaa9e7e00 [diffusion] fix: replace TextEncoderConfig with Qwen3TextConfig for Z-Image (#18560)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-02-12 14:31:23 +08:00
JooYeon
c9297661b9 fix: /metrics endpoint always reports engine_type="unified" in PD disaggregation mode (#18552)
Co-authored-by: joo_yeon.lee <joo_yeon.lee@samsung.com>
2026-02-12 14:20:43 +08:00
Li Jinliang
d91ce176bf Update README commands to include model-path option (#18557)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-02-12 14:15:26 +08:00
Zheng Li
4ed2548427 [Qwen3_5] Refactor Qwen3_5ForCausalLMMTP class implementation (#18538) 2026-02-12 13:38:26 +08:00
YAMY
454676811e [Flashinfer Autotune] Fix FlashInfer FP4 MoE autotuning crash by removing incorrect flatten on hidden_states_scale (#18500) 2026-02-12 13:31:27 +08:00
YC Tseng
20554a0a4f [AMD] rocm 7.2 image release, PR test, Nightly Test (#17799)
Co-authored-by: Alan Kao <akao@amd.com>
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com>
2026-02-11 21:29:25 -08:00
Ke Bao
93ede0db19 Update ci permission (#18693) 2026-02-12 13:25:43 +08:00
Liangsheng Yin
3404dda592 Try fix the max-parallel for maunally triggered test again. (#18686) 2026-02-11 20:08:25 -08:00
danielafrimi
e422bcaed8 [Mamba] Add float16 support for SSM cache dtype (#18444) 2026-02-12 11:27:47 +08:00
Zhiyu
7e262b6496 Update modelopt quantization config parsing (#13919)
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-02-12 11:08:29 +08:00
Liangsheng Yin
b2d7fd5c87 fix the max-parallel for /rerun-stage (#18658) 2026-02-11 19:06:51 -08:00
fy
123f57b84b update glm5 readme on npu (#18657) 2026-02-12 10:37:12 +08:00
Liangsheng Yin
10c6bee74f List more CI runs for pr-test (#18650) 2026-02-11 18:36:45 -08:00
R0CKSTAR
41e1fd0be7 [diffusion] fix: webui cannot correctly display generated video using wan2.2 (#18473)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-02-12 10:35:39 +08:00
liupeng374
c34832c02c glm5 md (#18655) 2026-02-12 10:11:59 +08:00
Jincong Chen
165aff38e1 Add CI permission for Chen-0210 (#18494) 2026-02-12 09:33:35 +08:00
Yi Zhong
dc1309fc7e Avoid kimi linear stream sync (#16186)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
2026-02-12 09:27:22 +08:00
Jiayi Yan
539bbf485c [Bugfix] fix config bug caused by PR #18273 (#18535) 2026-02-12 09:26:46 +08:00
Yuwei An
2bd8363486 [PCG] GPT OSS Triton Kernel Support (#18405)
Signed-off-by: Oasis-Git <ayw.sirius19@gmail.com>
2026-02-12 09:23:55 +08:00
qianyue76
f06ab17a73 [diffusion] docs: consolidate diffusion documentation into docs (#18095)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: JiaxinD <djx2048@gmail.com>
2026-02-11 16:55:07 -08:00
Alison Shao
7eaf866846 [CI] Install python3-dev for Triton JIT compilation on fresh runners (#18644) 2026-02-11 16:28:57 -08:00
Lianmin Zheng
5875ef0a34 Clean up noisy startup log messages and refactor loader.py (#18531) 2026-02-11 16:12:57 -08:00
Piotr Mazurek
ded068a76e Add LMF2 MoE model architecture (#17997) 2026-02-12 01:03:43 +08:00
Ke Bao
5d185efb78 Fix prefill stats for dllm (#18632) 2026-02-12 01:00:30 +08:00
Alison Shao
dcc63dc545 [CI] Guard python3 call in install script for fresh runners (#18609) 2026-02-12 00:05:29 +08:00
Vedant V Jhaveri
98b5013d59 add support to enable lora with embedding models (#17780)
Co-authored-by: Vedant Jhaveri <vjhaveri@linkedin.com>
2026-02-11 23:19:40 +08:00
Baizhou Zhang
947927bdb5 [V3.2] Change default CP token split method to --round-robin-split (#18613) 2026-02-11 20:14:35 +08:00
McZyWu
4f7422f7ba [NPU] support model skywork-reward-gemma2-2-27B-v0.2 (#16947)
Co-authored-by: cy <chenyang08056032@163.com>
2026-02-11 15:34:53 +08:00
sky
72c1526657 Register cp-atten-allgather buffers with symm memory (#17756)
Signed-off-by: wangfakang <fakangwang@gmail.com>
2026-02-11 15:26:37 +08:00
Thomas Wang
a8eef53dc4 Fp8 prefill attn kernel integration (#18528)
Co-authored-by: kkHuang-amd <wunhuang@amd.com>
2026-02-10 23:23:48 -08:00
BourneSun0527
2cc235e795 Fix Bug on dsv3.2 (#18553)
This PR affects only the NPU. If any issues arise, please contact iforgetmyname.
2026-02-11 14:39:01 +08:00
Michael
d84d2063d3 [AMD] Fix Janus-Pro crash and add Kimi-K2.5 nightly test (#18269) 2026-02-10 22:33:13 -08:00
Liangsheng Yin
cd90346a2b Add cache hit rate UT (#18566) 2026-02-10 21:27:41 -08:00
Liangsheng Yin
50f74285e9 Tiny fix regex warning (#18592) 2026-02-10 21:07:53 -08:00
cutetocute
8d2892330c chore: fix some typos (#18577)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2026-02-10 20:47:41 -08:00
赵晨阳
a2c38f7796 Enhance SMG guide with RL rollout systems benefits (#18588) 2026-02-10 20:20:45 -08:00
AlexZhao
3167bcc01c [Doc] Comprehensive Guide: Navigating DP, DPA, and SMG Best Practices (#18096)
Co-authored-by: 赵海源 <zhaohaiyuan@xiaohongshu.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>
2026-02-10 18:31:28 -08:00
Liangsheng Yin
93fca0bbc3 Fix wrong prefill log. (#18570) 2026-02-10 15:54:03 -08:00
Yi-Chia Chen
2bfab1bb67 Fix radix cache key to include generated tokens in multi-turn (regression) (#16521)
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
2026-02-10 14:08:34 -08:00
Thomas Wang
4262f5259b Tilelang sparse decode fwd for dsv32 mi355 (#18488)
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
2026-02-10 09:53:26 -08:00
Baizhou Zhang
2d38b8aca0 Revert "[sgl-kernel] upgrade deepgemm" (#18562) 2026-02-11 01:17:40 +08:00
husf
573ff55814 [NPU][docs]fix bug about hyperlink for best practice for ascend npu (#18561) 2026-02-10 20:03:28 +03:00
Zheng Li
44603764d6 fix(config): Support setting Mamba state dtype via config file (#18532)
Co-authored-by: 瑀澈 <yuche.lz@alibaba-inc.com>
2026-02-11 00:20:06 +08:00
Hexq0210
d0d387dea1 [NPU] update npu doc (#18474) 2026-02-10 16:59:13 +03:00
husf
99101ce30b [NPU][docs] improve docs for Best Practice on Ascend NPU (#18360) 2026-02-10 16:52:27 +03:00