This website requires JavaScript.
Explore
Help
Sign In
chenchenghao
/
sglang
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
Actions
12
Packages
Projects
Releases
Wiki
Activity
Files
25c7395934a92a213596d8bd9d00410207074796
sglang
/
sgl-kernel
/
python
/
sgl_kernel
T
History
Yineng Zhang
c5082f0f73
chore: fix cuda driver api issue and bump sgl-kernel 0.3.7.post1 (
#9746
)
2025-08-30 02:01:54 -07:00
..
testing
…
__init__.py
[NVIDA] [1/N] Nvfp4 Masked Gemm: Add quant op for the flashinfer grouped gemm (
#9200
)
2025-08-22 12:19:45 -07:00
allreduce.py
…
attention.py
…
cutlass_moe.py
…
elementwise.py
Add PDL support for quant kernel and rope kernel (
#9106
)
2025-08-20 01:56:29 -07:00
flash_attn.py
Reduce overhead for fa by not calling heavy CUDA property check (
#7375
)
2025-08-20 16:26:28 +08:00
fused_moe.py
…
gemm.py
[NVIDIA] [2/N] Optimize
silu_and_mul_scaled_fp4_grouped_quant
perf (
#9556
)
2025-08-29 17:17:03 -07:00
grammar.py
…
kvcacheio.py
[AMD] Support Hierarchical Caching on AMD GPUs (
#8236
)
2025-08-28 15:27:07 -07:00
marlin.py
…
memory.py
…
moe.py
…
sampling.py
…
scalar_type.py
…
sparse_flash_attn.py
…
spatial.py
…
speculative.py
…
top_k.py
…
utils.py
…
version.py
chore: fix cuda driver api issue and bump sgl-kernel 0.3.7.post1 (
#9746
)
2025-08-30 02:01:54 -07:00