Commit Graph
39 Commits
Author SHA1 Message Date
Артем Савкинandronnie_zheng 5297b02c88 [Diffusion] [NPU] Wan2.2-T2V-A14B-Diffusers modelslim quantization support (#17996)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-03-07 17:26:44 +03:00
liupeng374 5471e4a492 [NPU][Feature] eliminate dsv3 redundant rotary embed calculation (#19842) 2026-03-06 09:02:14 +08:00
KnightLTCandcy 1041f240c0 [NPU]grok2 model support (#17119)
Co-authored-by: cy <chenyang08056032@163.com>
2026-03-03 10:24:10 +08:00
JiaruiChang5268andJiaruiChang5268 b3718982a1 [Feature] add feature mla_ag_after_qlora for dsv3.2 (#19428)
Co-authored-by: JiaruiChang5268 <changjiarui1@huawei.com>
2026-03-02 20:00:31 +08:00
Michelle Wu b7f13a7b73 [NPU] bugs fix for Deepseek models (#19544) 2026-02-28 17:26:15 +08:00
Hexq0210andMcZyWu d0bb140034 [NPU] bugfix for model Qwen3-Coder-Next at weight shape transpose for npu. (#18700)
Co-authored-by: McZyWu <zhuoyun.wu.23@ucl.ac.uk>
2026-02-25 15:46:20 +08:00
fyandsglang-npu-bot 4f25a48d7a support xverse_moe on npu
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
2026-02-24 14:11:20 +08:00
SoluMilken 07a24f1a38 update pre-commit config (#18860) 2026-02-16 00:18:31 +08:00
McZyWuandcy 4f7422f7ba [NPU] support model skywork-reward-gemma2-2-27B-v0.2 (#16947)
Co-authored-by: cy <chenyang08056032@163.com>
2026-02-11 15:34:53 +08:00
Cheng Wan 84c09913eb Moving _alloc_extend_naive out of npu allocator (#18200) 2026-02-04 02:09:55 -08:00
khalilzhk b0a6d5244c [NPU] support dsv32 radixcache on ascend (#17964) 2026-02-03 03:34:12 +08:00
jiashaokun-1 0fe282543f [NPU] support the Enable return routed experts (#17025) 2026-02-01 18:31:39 +08:00
e86476acfc [NPU] support llama-3.2-11B-vision-instruct mode for NPU (#17492)
Co-authored-by: McZyWu <zhuoyun.wu.23@ucl.ac.uk>
Co-authored-by: chenyang08056032 <chenyang08056032@163.com>
Co-authored-by: Hexq0210 <893781835@qq.com>
2026-01-31 08:49:38 +08:00
McZyWuandcy 70db3398d1 [NPU] enhance accuracy for model kimi-vl-a3b-instruct (#17480)
Co-authored-by: cy <chenyang08056032@163.com>
2026-01-30 15:19:42 +08:00
Артем Савкинandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> b77b0ffd60 [NPU] NZ for non-quantized MOE, Qwen3 MOE double memory consumption fix (#15904)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-29 00:55:08 +08:00
monkeyLovedingandKelon d578b41bad [NPU] Adapt cann 8.5: use sfa and lightning indexer op from cann and CI update (#17615)
Co-authored-by: Kelon <kelonlu@163.com>
2026-01-27 19:03:53 +08:00
b56366f827 [NPU]DeepSeek-V3.2 support npu mlaprolog (#15381)
Co-authored-by: Zhengda Qin <zhengdqin@gmail.com>
Co-authored-by: richhuan <huan_rz@qq.com>
2026-01-26 20:42:37 +08:00
McZyWuandcy 2734b23481 accuracy enhancement for baichuan2-13B for npu (#16868)
Co-authored-by: cy <chenyang08056032@163.com>
2026-01-26 16:14:35 +08:00
McZyWu 8a5ed2434f [NPU]support model MiniCPM3-4B for npu (#16866) 2026-01-24 08:25:12 +08:00
c0b5a180fe [NPU]bugfix: fix for dsv3.2 and dsvl2 (#17007)
Co-authored-by: Hexq0210 <893781835@qq.com>
Co-authored-by: liupeng374 <782420244@qq.com>
Co-authored-by: cy <chenyang08056032@163.com>
2026-01-23 11:15:15 +08:00
5d299c25c0 [NPU] bugfix with Kimi-k2 and bge-reranker-v2 model (#17478)
Co-authored-by: amote-i <49533125+amote-i@users.noreply.github.com>
Co-authored-by: cy <chenyang08056032@163.com>
2026-01-22 22:02:05 +08:00
a95c9f5b81 [NPU] Remove paged attention & Change fia to default attention (#17394)
Co-authored-by: Liwansi <62291011+Liwansi@users.noreply.github.com>
Co-authored-by: chenxu214 <justin_cc2025@163.com>
Co-authored-by: chenyang08056032 <chenyang08056032@163.com>
2026-01-21 23:58:25 +08:00
Todobe 733de6be31 [NPU]Support GPT-OSS for NPU (#14197) 2026-01-19 04:13:41 +08:00
424a380077 [NPU] NPU quantization refactoring & more quantization formats support (#14504)
Co-authored-by: TamirBaydasov <mr.jeijy@gmail.com>
Co-authored-by: Tamir Baydasov <41994229+TamirBaydasov@users.noreply.github.com>
Co-authored-by: Савкин Артем <savkinartem@MacBook-Air-Viktoria.local>
Co-authored-by: Edward Shogulin <edward.shogulin@gmail.com>
2026-01-15 04:25:15 +08:00
hw-csong 261860e17b [NPU][Bugfix] move free_page logics to cpu (#16608) 2026-01-08 12:00:18 +08:00
Yongfei Xu 0d244116d2 [DeepSeek v3.2] opt Context Parallelism: support fused moe, multi batch and fp8 kvcache (#13959) 2026-01-02 23:49:14 +08:00
jiaming1130andZhengdQin 60a230b1fd [NPU] Support w4a8 with activation clip (#14736)
Co-authored-by: ZhengdQin <46387172+ZhengdQin@users.noreply.github.com>
2025-12-27 16:19:46 +08:00
Hexq0210 3fd232ad96 [NPU] update Mixed chunk op to FIA (#15518) 2025-12-26 12:17:40 +08:00
Артем Савкин 3e01f3a533 [NPU] [BUGFIX] [CRITICAL!] Fix NPU inference (torch_npu._npu_reshape_and_cache() crash) (#15484) 2025-12-20 14:53:46 +08:00
Hexq0210 241ae17b25 [NPU] bugfix for chunkedprefill (#15166) 2025-12-19 22:41:17 +08:00
Even Zhou 71cb90378b [NPU] fix for NPU memory settings logic (#15258) 2025-12-16 17:04:22 -08:00
XDaoHongandZhengdQin 4733fcff1f [Feature] npu support enable_torch_compile for torchair backend (#13410)
Co-authored-by: ZhengdQin <zhengdqin@gmail.com>
2025-12-16 09:23:51 +08:00
Liwansi 30da2f0598 [NPU][eagle3] support qwen eagle3 on NPU (#14820) 2025-12-16 02:25:13 +08:00
liupeng374 d36299ad77 [NPU] perf update with kvcache nz & w4a8 quant (#14423) 2025-12-13 17:39:55 +08:00
ZhengdQin c05d3afb5d [NPU] optimization for dsv3.2 (#14572) 2025-12-12 14:52:16 +08:00
liupeng374 388018a5bd [NPU] adapt dsv3.2 nsa prefill context parallel (#14541) 2025-12-11 20:12:59 +08:00
liupeng374 b8cfa02c01 [NPU] bug fix for mtp and w4a8 (#14806) 2025-12-10 22:55:13 +08:00
khalilzhk 948b6acee8 [BugFix] fix prefixcache performance and accuracy on ascend (#13573) 2025-12-08 02:16:20 +08:00
Even Zhou 894c0dc57c [NPU][1/N] NPU basic functions refactor and new modelslim quant type (#13359) 2025-12-04 16:15:31 +08:00