This website requires JavaScript.
Explore
Help
Sign In
chenchenghao
/
sglang
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
Actions
12
Packages
Projects
Releases
Wiki
Activity
Files
451ffe74d9071a2f67194e2497152643b0b809b0
sglang
/
sgl-kernel
/
python
/
sgl_kernel
T
History
JieXin Liang
18efb5e8e0
[perf][sgl-kernel] extend cutlass_mla_decode to support num_head < 128 (
#6929
)
2025-06-08 19:37:34 -07:00
..
__init__.py
Add a CUDA kernel for fusing mapping and weighted sum for MoE. (
#6916
)
2025-06-07 15:24:39 -07:00
allreduce.py
support 1 shot allreduce in 1-node and 2-node using mscclpp (
#6277
)
2025-06-04 22:11:24 -07:00
attention.py
[perf][sgl-kernel] extend cutlass_mla_decode to support num_head < 128 (
#6929
)
2025-06-08 19:37:34 -07:00
elementwise.py
…
flash_attn.py
…
gemm.py
[1/2] Add Kernel support for Cutlass based Fused FP4 MoE (
#6093
)
2025-06-02 13:48:03 -07:00
grammar.py
…
moe.py
Add a CUDA kernel for fusing mapping and weighted sum for MoE. (
#6916
)
2025-06-07 15:24:39 -07:00
sampling.py
…
sparse_flash_attn.py
…
speculative.py
…
utils.py
misc: cache is_hopper_arch (
#6799
)
2025-06-01 15:28:31 -07:00
version.py
chore: bump sgl-kernel v0.1.7 (
#6963
)
2025-06-08 02:43:15 -07:00