This website requires JavaScript.
Explore
Help
Sign In
chenchenghao
/
sglang
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
Actions
11
Packages
Projects
Releases
Wiki
Activity
Files
76eb1c8406bc4a73cec05f516093ef80d8e56387
sglang
/
sgl-kernel
/
python
/
sgl_kernel
T
History
3 people
jianan-gu
Ma Mingfei
Fan Yin
336dc4579e
[CPU] Optimize Qwen3-next model on CPU (
#12525
)
...
Co-authored-by: Ma Mingfei <
mingfei.ma@intel.com
> Co-authored-by: Fan Yin <
1106310035@qq.com
>
2026-01-29 22:03:58 -08:00
..
quantization
…
testing
[1/2] Add rope kernel in sgl-kernel (
#14334
)
2025-12-04 16:45:44 +08:00
__init__.py
[CPU] Optimize Qwen3-next model on CPU (
#12525
)
2026-01-29 22:03:58 -08:00
_fa4_interface.py
[NVIDIA] upstream FA4 (
#15182
)
2026-01-11 15:31:28 +08:00
allreduce.py
[amd] Add deterministic all-reduce kernel for AMD (ROCm) (
#15340
)
2025-12-18 23:36:03 -08:00
attention.py
…
cutlass_moe.py
…
elementwise.py
[Diffusion] Delete sgl-kernel outdated time_embedding kernel (
#17278
)
2026-01-28 14:18:53 +08:00
expert_specialization.py
[sgl-kernel][Feat][B200][1/N] Support MXFP8 Grouped GEMM in Blackwell (
#13731
)
2025-12-04 10:09:09 +08:00
flash_attn.py
[NVIDIA] upstream FA4 (
#15182
)
2026-01-11 15:31:28 +08:00
flash_mla.py
…
fused_moe.py
Add new moe wna16 marlin gemm (
#14122
)
2025-12-01 23:07:53 +08:00
gemm.py
…
grammar.py
…
hadamard.py
…
kvcacheio.py
…
load_utils.py
[sgl-kernel] fix runtime error while preloading CUDA runtime (
#13089
)
2025-12-02 23:03:51 +08:00
mamba.py
[CPU] Optimize Qwen3-next model on CPU (
#12525
)
2026-01-29 22:03:58 -08:00
marlin.py
…
memory.py
…
moe.py
[sgl-kernel][1/2] Fused qk_norm_rope for GLM4.6 (
#15141
)
2025-12-18 17:07:04 +08:00
sampling.py
…
scalar_type.py
…
sparse_flash_attn.py
…
spatial.py
…
speculative.py
…
test_utils.py
…
top_k.py
…
utils.py
…
version.py
chore: bump sgl-kernel version to 0.3.21 (
#16888
)
2026-01-14 13:27:49 +08:00