[Diffusion] [NPU] Wan2.2-T2V-A14B-Diffusers modelslim quantization support (#17996)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
This commit is contained in:
co-authored by
ronnie_zheng
parent
f8d4eb7022
commit
5297b02c88
@@ -19,3 +19,8 @@ Compressed-tensors (LLM Compressor) on Ascend support:
|
||||
- [x] [W4A16 MOE](https://github.com/sgl-project/sglang/pull/12759)
|
||||
- [x] [W8A8 dynamic linear](https://github.com/sgl-project/sglang/pull/14504)
|
||||
- [x] [W8A8 dynamic MOE](https://github.com/sgl-project/sglang/pull/14504)
|
||||
|
||||
Diffusion model [modelslim](https://github.com/sgl-project/sglang/pull/17996) quantization on Ascend support:
|
||||
- [x] W4A4 dynamic linear
|
||||
- [x] W8A8 static linear
|
||||
- [x] W8A8 dynamic linear
|
||||
|
||||
Reference in New Issue
Block a user