 Артем Савкинandronnie_zheng
|
5297b02c88
|
[Diffusion] [NPU] Wan2.2-T2V-A14B-Diffusers modelslim quantization support (#17996)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-03-07 17:26:44 +03:00 |
|
liupeng374
|
5471e4a492
|
[NPU][Feature] eliminate dsv3 redundant rotary embed calculation (#19842)
|
2026-03-06 09:02:14 +08:00 |
|
 KnightLTCandcy
|
1041f240c0
|
[NPU]grok2 model support (#17119)
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-03-03 10:24:10 +08:00 |
|
 JiaruiChang5268andJiaruiChang5268
|
b3718982a1
|
[Feature] add feature mla_ag_after_qlora for dsv3.2 (#19428)
Co-authored-by: JiaruiChang5268 <changjiarui1@huawei.com>
|
2026-03-02 20:00:31 +08:00 |
|
Michelle Wu
|
b7f13a7b73
|
[NPU] bugs fix for Deepseek models (#19544)
|
2026-02-28 17:26:15 +08:00 |
|
 Hexq0210andMcZyWu
|
d0bb140034
|
[NPU] bugfix for model Qwen3-Coder-Next at weight shape transpose for npu. (#18700)
Co-authored-by: McZyWu <zhuoyun.wu.23@ucl.ac.uk>
|
2026-02-25 15:46:20 +08:00 |
|
 fyandsglang-npu-bot
|
4f25a48d7a
|
support xverse_moe on npu
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
|
2026-02-24 14:11:20 +08:00 |
|
SoluMilken
|
07a24f1a38
|
update pre-commit config (#18860)
|
2026-02-16 00:18:31 +08:00 |
|
 McZyWuandcy
|
4f7422f7ba
|
[NPU] support model skywork-reward-gemma2-2-27B-v0.2 (#16947)
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-02-11 15:34:53 +08:00 |
|
Cheng Wan
|
84c09913eb
|
Moving _alloc_extend_naive out of npu allocator (#18200)
|
2026-02-04 02:09:55 -08:00 |
|
khalilzhk
|
b0a6d5244c
|
[NPU] support dsv32 radixcache on ascend (#17964)
|
2026-02-03 03:34:12 +08:00 |
|
jiashaokun-1
|
0fe282543f
|
[NPU] support the Enable return routed experts (#17025)
|
2026-02-01 18:31:39 +08:00 |
|
  
|
e86476acfc
|
[NPU] support llama-3.2-11B-vision-instruct mode for NPU (#17492)
Co-authored-by: McZyWu <zhuoyun.wu.23@ucl.ac.uk>
Co-authored-by: chenyang08056032 <chenyang08056032@163.com>
Co-authored-by: Hexq0210 <893781835@qq.com>
|
2026-01-31 08:49:38 +08:00 |
|
 McZyWuandcy
|
70db3398d1
|
[NPU] enhance accuracy for model kimi-vl-a3b-instruct (#17480)
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-01-30 15:19:42 +08:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) Артем Савкинandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
b77b0ffd60
|
[NPU] NZ for non-quantized MOE, Qwen3 MOE double memory consumption fix (#15904)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-01-29 00:55:08 +08:00 |
|
 monkeyLovedingandKelon
|
d578b41bad
|
[NPU] Adapt cann 8.5: use sfa and lightning indexer op from cann and CI update (#17615)
Co-authored-by: Kelon <kelonlu@163.com>
|
2026-01-27 19:03:53 +08:00 |
|
 
|
b56366f827
|
[NPU]DeepSeek-V3.2 support npu mlaprolog (#15381)
Co-authored-by: Zhengda Qin <zhengdqin@gmail.com>
Co-authored-by: richhuan <huan_rz@qq.com>
|
2026-01-26 20:42:37 +08:00 |
|
 McZyWuandcy
|
2734b23481
|
accuracy enhancement for baichuan2-13B for npu (#16868)
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-01-26 16:14:35 +08:00 |
|
McZyWu
|
8a5ed2434f
|
[NPU]support model MiniCPM3-4B for npu (#16866)
|
2026-01-24 08:25:12 +08:00 |
|
  
|
c0b5a180fe
|
[NPU]bugfix: fix for dsv3.2 and dsvl2 (#17007)
Co-authored-by: Hexq0210 <893781835@qq.com>
Co-authored-by: liupeng374 <782420244@qq.com>
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-01-23 11:15:15 +08:00 |
|
 
|
5d299c25c0
|
[NPU] bugfix with Kimi-k2 and bge-reranker-v2 model (#17478)
Co-authored-by: amote-i <49533125+amote-i@users.noreply.github.com>
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-01-22 22:02:05 +08:00 |
|
  
|
a95c9f5b81
|
[NPU] Remove paged attention & Change fia to default attention (#17394)
Co-authored-by: Liwansi <62291011+Liwansi@users.noreply.github.com>
Co-authored-by: chenxu214 <justin_cc2025@163.com>
Co-authored-by: chenyang08056032 <chenyang08056032@163.com>
|
2026-01-21 23:58:25 +08:00 |
|
Todobe
|
733de6be31
|
[NPU]Support GPT-OSS for NPU (#14197)
|
2026-01-19 04:13:41 +08:00 |
|
   
|
424a380077
|
[NPU] NPU quantization refactoring & more quantization formats support (#14504)
Co-authored-by: TamirBaydasov <mr.jeijy@gmail.com>
Co-authored-by: Tamir Baydasov <41994229+TamirBaydasov@users.noreply.github.com>
Co-authored-by: Савкин Артем <savkinartem@MacBook-Air-Viktoria.local>
Co-authored-by: Edward Shogulin <edward.shogulin@gmail.com>
|
2026-01-15 04:25:15 +08:00 |
|
hw-csong
|
261860e17b
|
[NPU][Bugfix] move free_page logics to cpu (#16608)
|
2026-01-08 12:00:18 +08:00 |
|
Yongfei Xu
|
0d244116d2
|
[DeepSeek v3.2] opt Context Parallelism: support fused moe, multi batch and fp8 kvcache (#13959)
|
2026-01-02 23:49:14 +08:00 |
|
 jiaming1130andZhengdQin
|
60a230b1fd
|
[NPU] Support w4a8 with activation clip (#14736)
Co-authored-by: ZhengdQin <46387172+ZhengdQin@users.noreply.github.com>
|
2025-12-27 16:19:46 +08:00 |
|
Hexq0210
|
3fd232ad96
|
[NPU] update Mixed chunk op to FIA (#15518)
|
2025-12-26 12:17:40 +08:00 |
|
Артем Савкин
|
3e01f3a533
|
[NPU] [BUGFIX] [CRITICAL!] Fix NPU inference (torch_npu._npu_reshape_and_cache() crash) (#15484)
|
2025-12-20 14:53:46 +08:00 |
|
Hexq0210
|
241ae17b25
|
[NPU] bugfix for chunkedprefill (#15166)
|
2025-12-19 22:41:17 +08:00 |
|
Even Zhou
|
71cb90378b
|
[NPU] fix for NPU memory settings logic (#15258)
|
2025-12-16 17:04:22 -08:00 |
|
 XDaoHongandZhengdQin
|
4733fcff1f
|
[Feature] npu support enable_torch_compile for torchair backend (#13410)
Co-authored-by: ZhengdQin <zhengdqin@gmail.com>
|
2025-12-16 09:23:51 +08:00 |
|
Liwansi
|
30da2f0598
|
[NPU][eagle3] support qwen eagle3 on NPU (#14820)
|
2025-12-16 02:25:13 +08:00 |
|
liupeng374
|
d36299ad77
|
[NPU] perf update with kvcache nz & w4a8 quant (#14423)
|
2025-12-13 17:39:55 +08:00 |
|
ZhengdQin
|
c05d3afb5d
|
[NPU] optimization for dsv3.2 (#14572)
|
2025-12-12 14:52:16 +08:00 |
|
liupeng374
|
388018a5bd
|
[NPU] adapt dsv3.2 nsa prefill context parallel (#14541)
|
2025-12-11 20:12:59 +08:00 |
|
liupeng374
|
b8cfa02c01
|
[NPU] bug fix for mtp and w4a8 (#14806)
|
2025-12-10 22:55:13 +08:00 |
|
khalilzhk
|
948b6acee8
|
[BugFix] fix prefixcache performance and accuracy on ascend (#13573)
|
2025-12-08 02:16:20 +08:00 |
|
Even Zhou
|
894c0dc57c
|
[NPU][1/N] NPU basic functions refactor and new modelslim quant type (#13359)
|
2025-12-04 16:15:31 +08:00 |
|