Logo
Explore Help
Sign In
chenchenghao/sglang
1
0
Fork 0
You've already forked sglang
Code Issues Pull Requests Actions 12 Packages Projects Releases Wiki Activity
Files
57eec0bfbce964e347ef2affb999e03416f22325
sglang/python/sglang/srt/layers/attention
History
lukec 57eec0bfbc fix FlashMLA cudagraph config (#4691)
Co-authored-by: yinfan98 <1106310035@qq.com>
2025-03-24 21:06:58 -07:00
..
triton_ops
[fix] fix illegal mem access and clean up triton attention backend (#4571)
2025-03-20 02:01:52 -07:00
base_attn_backend.py
[fix] fix illegal mem access and clean up triton attention backend (#4571)
2025-03-20 02:01:52 -07:00
double_sparsity_backend.py
Misc clean up; Remove the support of jump forward (#4032)
2025-03-03 07:02:14 -08:00
flashattention_backend.py
Support FA3 as Attention backend by using --attention-backend fa3 (#4680)
2025-03-23 23:28:11 -07:00
flashinfer_backend.py
[fix] fix illegal mem access and clean up triton attention backend (#4571)
2025-03-20 02:01:52 -07:00
flashinfer_mla_backend.py
[fix] fix illegal mem access and clean up triton attention backend (#4571)
2025-03-20 02:01:52 -07:00
flashmla_backend.py
fix FlashMLA cudagraph config (#4691)
2025-03-24 21:06:58 -07:00
torch_native_backend.py
Misc clean up; Remove the support of jump forward (#4032)
2025-03-03 07:02:14 -08:00
triton_backend.py
[fix] fix illegal mem access and clean up triton attention backend (#4571)
2025-03-20 02:01:52 -07:00
utils.py
Support FlashMLA backend cuda graph (#4514)
2025-03-19 08:25:34 -07:00
vision.py
refactor: bug fixes and refactor for vlm (#4661)
2025-03-22 22:48:49 -07:00
Powered by Gitea Version: 1.25.5 Page: 62ms Template: 2ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API