This website requires JavaScript.
Explore
Help
Sign In
chenchenghao
/
sglang
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
Actions
12
Packages
Projects
Releases
Wiki
Activity
Files
2f66b0671b4204d811c828b6e2331a4361fe637b
sglang
/
python
/
sglang
/
srt
/
layers
/
attention
/
nsa
T
History
Hudson Xing
9d878c1f3e
Optimize FP8 MLA KV cache writes with Triton kernel (
#15522
)
2025-12-25 12:35:39 -08:00
..
dequant_k_cache.py
[NSA] Fix NSA backend assertion error when running DeepSeek-V3.2 PP with radix-cache (
#15086
)
2025-12-14 17:13:18 -08:00
index_buf_accessor.py
[Performance] Optimize NSA Indexer K/S Buffer Access with Fused Triton Kernels (
#13812
)
2025-12-02 18:53:06 -08:00
nsa_backend_mtp_precompute.py
[Performance] optimize NSA backend metadata computation for multi-step speculative decoding (
#14781
)
2025-12-18 13:48:27 -08:00
nsa_indexer.py
[DSv32] Move deep_gemm.get_paged_mqa_logits_metadata to init time as metadata (
#15040
)
2025-12-19 13:23:33 -08:00
quant_k_cache.py
Optimize FP8 MLA KV cache writes with Triton kernel (
#15522
)
2025-12-25 12:35:39 -08:00
tilelang_kernel.py
…
transform_index.py
[DeepSeekV32] Bug fix to ensure
page_table
and
result
in same type (
#12300
)
2025-10-30 18:56:51 -07:00
triton_kernel.py
[DPSKv3.2] Rewrite nsa tilelang act_quant kernel to triton (
#11450
)
2025-10-10 23:13:46 -07:00
utils.py
[Performance] optimize NSA backend metadata computation for multi-step speculative decoding (
#14781
)
2025-12-18 13:48:27 -08:00