[NPU] NPU quantization refactoring & more quantization formats support (#14504)
Co-authored-by: TamirBaydasov <mr.jeijy@gmail.com> Co-authored-by: Tamir Baydasov <41994229+TamirBaydasov@users.noreply.github.com> Co-authored-by: Савкин Артем <savkinartem@MacBook-Air-Viktoria.local> Co-authored-by: Edward Shogulin <edward.shogulin@gmail.com>
This commit is contained in:
@@ -30,7 +30,6 @@ python3 -m sglang.launch_server \
|
||||
--trust-remote-code \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--quantization modelslim \
|
||||
--watchdog-timeout 9000 \
|
||||
--cuda-graph-bs 8 16 24 28 32 \
|
||||
--mem-fraction-static 0.68 \
|
||||
@@ -88,7 +87,6 @@ python -m sglang.launch_server \
|
||||
--mem-fraction-static 0.6 \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--quantization modelslim \
|
||||
--max-running-requests 8 \
|
||||
--context-length 8192 \
|
||||
--disable-radix-cache \
|
||||
@@ -144,7 +142,6 @@ python -m sglang.launch_server \
|
||||
--max-running-requests 352 \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--quantization modelslim \
|
||||
--moe-a2a-backend deepep \
|
||||
--enable-dp-attention \
|
||||
--deepep-mode low_latency \
|
||||
@@ -216,7 +213,6 @@ do
|
||||
--mem-fraction-static 0.81 \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--quantization modelslim \
|
||||
--max-running-requests 8 \
|
||||
--context-length 8192 \
|
||||
--disable-radix-cache \
|
||||
@@ -279,7 +275,6 @@ do
|
||||
--max-running-requests 832 \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--quantization modelslim \
|
||||
--moe-a2a-backend deepep \
|
||||
--enable-dp-attention \
|
||||
--deepep-mode low_latency \
|
||||
|
||||
21
docs/platforms/ascend_npu_quantization.md
Normal file
21
docs/platforms/ascend_npu_quantization.md
Normal file
@@ -0,0 +1,21 @@
|
||||
Quantization on Ascend.
|
||||
|
||||
To load already quantized models, simply load the model weights and config. Again, if the model has been quantized offline, there's no need to add `--quantization` argument when starting the engine. The quantization method will be automatically parsed from the downloaded `quant_model_description.json` or `config.json` config.
|
||||
|
||||
[ModelSlim on Ascend support](https://github.com/sgl-project/sglang/pull/14504):
|
||||
- [x] W4A4 dynamic linear
|
||||
- [x] W8A8 static linear
|
||||
- [x] W8A8 dynamic linear
|
||||
- [x] W4A8 dynamic MOE
|
||||
- [x] W8A8 dynamic MOE
|
||||
|
||||
[AWQ on Ascend support](https://github.com/sgl-project/sglang/pull/10158):
|
||||
- [x] W4A16 linear
|
||||
- [x] W8A16 linear # Need to test
|
||||
- [x] W4A16 MOE # Need to test
|
||||
|
||||
Compressed-tensors (LLM Compressor) on Ascend support:
|
||||
- [x] [W4A8 dynamic MOE with/without activation clip](https://github.com/sgl-project/sglang/pull/14736) # Need to test
|
||||
- [x] [W4A16 MOE](https://github.com/sgl-project/sglang/pull/12759)
|
||||
- [x] [W8A8 dynamic linear](https://github.com/sgl-project/sglang/pull/14504)
|
||||
- [x] [W8A8 dynamic MOE](https://github.com/sgl-project/sglang/pull/14504)
|
||||
Reference in New Issue
Block a user