This website requires JavaScript.
Explore
Help
Sign In
chenchenghao
/
sglang
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
Actions
11
Packages
Projects
Releases
Wiki
Activity
Files
18c27131f5515f5ef12742bb2e06b3e0345c0340
sglang
/
sgl-kernel
/
benchmark
T
History
Qingquan Song
4068e01292
Fix per token fp8 quant precision (
#4362
)
2025-03-12 21:19:05 -07:00
..
bench_awq_dequant.py
Add awq dequantize kernel to sgl with 1x to 3x speedup (
#4104
)
2025-03-12 00:10:02 -07:00
bench_cublas_grouped_gemm.py
[Feature] Apply Cublas Grouped Gemm kernel (
#3629
)
2025-02-18 15:18:31 +08:00
bench_fp8_blockwise_gemm.py
support blockwise fp8 matmul kernel (
#3267
)
2025-02-13 01:49:33 +08:00
bench_fp8_gemm.py
support w8a8 fp8 kernel with CUTLASS (
#3047
)
2025-01-26 15:46:51 +08:00
bench_int8_gemm.py
…
bench_lightning_attention_decode.py
…
bench_moe_align_block_size.py
feat: support ep size < 32 for sgl kernel (
#4348
)
2025-03-12 20:50:46 -07:00
bench_per_tensor_quant_fp8.py
[quant kernel] sgl-kernel support per_tensor_quant fp8 (
#3786
)
2025-03-06 18:05:43 -08:00
bench_per_token_group_quant_fp8.py
[Refactor] Reducing code duplication across FP8 CUDA quantization kernels (
#4163
)
2025-03-06 22:58:52 -08:00
bench_per_token_quant_fp8.py
Fix per token fp8 quant precision (
#4362
)
2025-03-12 21:19:05 -07:00