Ke Bao
|
fb683be6eb
|
Use attn tp group in embedding for more models (#17570)
|
2026-01-24 13:37:44 +08:00 |
|
DarkSharpness
|
8e43980ebb
|
[Feature] JIT Fused QK norm + qk norm clean up (#15835)
|
2025-12-28 11:53:50 +08:00 |
|
Yuhao Yao
|
e9e7f15eb5
|
[bugfix] fix TBO crashes when attn_tp_size > 1 (#13730)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2025-12-11 18:18:40 -08:00 |
|
b8zhong
|
6d5d76ad97
|
remove unecessary dual stream token threshold from the rest of models (qwen moe, kimi linear, etc.) (#14337)
|
2025-12-06 19:57:26 -08:00 |
|
Zehuan Li
|
21b0582d4b
|
[feature] Initial block diffusion language model support (#12588)
Co-authored-by: Tiwei Bie <tiwei.btw@antgroup.com>
|
2025-11-26 17:57:54 +08:00 |
|