Logo
Explore Help
Sign In
chenchenghao/sglang
Watch 1
Star 0
Fork 0
Code Issues Pull Requests Actions 12 Packages Projects Releases Wiki Activity
2,236 Commits 1 Branch 0 Tags
d3d4d76758b15c2c03e37e82cb85044f45332bfa
Commit Graph
10 Commits
Author SHA1 Message Date
Lianmin Zheng 935cda944b Misc clean up; Remove the support of jump forward (#4032) 2025-03-03 07:02:14 -08:00
Ke Bao de5533341e Update Triton extend backend interface (#3309) 2025-02-05 18:12:22 +08:00
Lianmin Zhengandyukavio f44d143949 Support target model verification in the attention backend (#2678)
Co-authored-by: yukavio <kavioyu@gmail.com>
2024-12-30 22:58:55 -08:00
Ke Bao ec52464dde MLA prefill w/o weight absorption (#2349) 2024-12-05 01:50:28 +08:00
Lianmin Zheng 384d85ba35 Re-introduce get_cuda_graph_seq_len_fill_value (#1783) 2024-10-24 13:30:11 -07:00
Lianmin Zheng fc82f5a743 [Fix] Fix cuda graph padding for triton attention backend (#1782) 2024-10-24 12:33:15 -07:00
Liangsheng Yin 94cde10920 Llama3.2 vision model support (#1551) 2024-10-21 15:01:21 -07:00
Lianmin Zheng 09603c6dc9 Maintain seq_lens_sum to make more FlashInfer operations non-blocking (#1741) 2024-10-21 01:43:16 -07:00
Lianmin Zheng 6d0fa73ece Simplify flashinfer utilities (#1704) 2024-10-17 22:54:14 -07:00
Shuo Yang 061e546313 Support double sparsity (#1459) 2024-10-14 02:00:41 -07:00
Powered by Gitea Version: 1.27.2 Page: 861ms Template: 2ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API