Logo
Explore Help
Sign In
chenchenghao/sglang
Watch 1
Star 0
Fork 0
Code Issues Pull Requests Actions 12 Packages Projects Releases Wiki Activity
Files
611720919d0f6cf8a481f19bbd4046dcba9a9130
sglang/python/sglang/srt/layers/quantization
T
History
Yineng Zhang 611720919d fix: use deepgemm only on hopper (#5310)
2025-04-11 20:48:24 -07:00
..
compressed_tensors
Support Llama4 fp8 inference (#5194)
2025-04-09 20:14:34 +08:00
configs
…
__init__.py
FP4 weight loading and inference (2/2) (#3972)
2025-04-08 17:26:21 -07:00
awq.py
Clean up import vllm in quantization/__init__.py (#4834)
2025-03-28 10:34:10 -07:00
base_config.py
…
blockwise_int8.py
Add Llama4 support (#5092)
2025-04-07 00:29:36 -07:00
fp8_kernel.py
fix: use deepgemm only on hopper (#5310)
2025-04-11 20:48:24 -07:00
fp8_utils.py
Support Llama4 fp8 inference (#5194)
2025-04-09 20:14:34 +08:00
fp8.py
ROCm/AITER CK_MoE: update 2-stage kernels & support both Activations (#5228)
2025-04-10 18:19:57 -07:00
gptq.py
Clean up import vllm in quantization/__init__.py (#4834)
2025-03-28 10:34:10 -07:00
int8_kernel.py
…
int8_utils.py
…
kv_cache.py
Fix loading KV quantization scale; Enable modelopt kv cache (#4686)
2025-04-08 09:11:35 -07:00
modelopt_quant.py
FP4 weight loading and inference (2/2) (#3972)
2025-04-08 17:26:21 -07:00
moe_wna16.py
Add Llama4 support (#5092)
2025-04-07 00:29:36 -07:00
utils.py
[2/3] fix dsv3 awq issue (#4625)
2025-04-03 17:36:39 -07:00
w8a8_fp8.py
Support Llama4 fp8 inference (#5194)
2025-04-09 20:14:34 +08:00
w8a8_int8.py
Support Llama4 fp8 inference (#5194)
2025-04-09 20:14:34 +08:00
Powered by Gitea Version: 1.27.2 Page: 2645ms Template: 89ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API