Minglei Zhu
|
d90c0837e5
|
[hybrid-model] clean up and consolidate redundant fields in RadixLinearAttention (#17660)
|
2026-01-27 10:37:58 -08:00 |
|
Yuan Luo
|
1e8db18290
|
[Kimi-Linear] Remove duplicated code in kimi-linear (#17731)
|
2026-01-26 14:20:24 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
0c8165ffbd
|
[Kimi-Linear] Refactor Kimi-Linear to support RadixLinearAttention (#17506)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-01-24 21:27:13 +08:00 |
|
strgrb
|
176da1bbdd
|
Fix: mistake sigmoid in kda (#17508)
|
2026-01-24 13:35:14 +08:00 |
|
strgrb
|
bcc6d84f93
|
Use fused_sigmoid_gating_delta_rule_update_kernel for KDA (#17108)
|
2026-01-21 19:24:29 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
e6b7c04947
|
[Kimi-Linear] Refactor kimi-linear gate calculation to avoid duplicated code (#17160)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-01-20 14:29:24 +08:00 |
|
b8zhong
|
6d5d76ad97
|
remove unecessary dual stream token threshold from the rest of models (qwen moe, kimi linear, etc.) (#14337)
|
2025-12-06 19:57:26 -08:00 |
|
b8zhong
|
cc2e36c352
|
overlap shared + routed expert computation in kimi linear (#12660)
|
2025-11-11 14:52:58 -08:00 |
|
 Ke Baoandyizhang2077
|
a4bf5c6ad2
|
Support Kimi Linear (#12469)
Co-authored-by: yizhang2077 <1109276519@qq.com>
|
2025-10-31 14:03:35 -07:00 |
|