Logo
Explore Help
Sign In
chenchenghao/sglang
Watch 1
Star 0
Fork 0
Code Issues Pull Requests Actions 11 Packages Projects Releases Wiki Activity
Files
014cab4dd23eb637f18d887f767a1791db1ae8ac
sglang/python/sglang/srt/layers
T
History
Yineng Zhang 014cab4dd2 update forward_return_lse (#3425)
2025-02-09 20:18:44 +08:00
..
attention
update forward_return_lse (#3425)
2025-02-09 20:18:44 +08:00
moe
optimize moe_align_kernel cuda (#3347)
2025-02-07 00:53:46 +08:00
quantization
Add H20 fp8 w8a8 gemm config (#3386)
2025-02-08 15:36:31 +08:00
activation.py
support QuickGELU (#3250)
2025-02-01 19:31:47 +08:00
dp_attention.py
Return more infos for computing average acceptance length (#3152)
2025-01-26 04:51:54 -08:00
layernorm.py
update and simplify CustomOp (#3249)
2025-02-01 18:56:44 +08:00
linear.py
Improve weight loading and code style (#3174)
2025-01-27 03:00:41 -08:00
logits_processor.py
Use int64 as indices for set_kv_buffer (#3039)
2025-01-21 19:46:09 -08:00
parameter.py
Improve weight loading and code style (#3174)
2025-01-27 03:00:41 -08:00
pooler.py
Rename InputMetadata -> ForwardBatch (#1543)
2024-09-30 02:41:11 -07:00
radix_attention.py
Fix perf regression on small batch sizes (#3008)
2025-01-20 03:39:49 -08:00
rotary_embedding.py
update and simplify CustomOp (#3249)
2025-02-01 18:56:44 +08:00
sampler.py
Bug: Fix min_p sampling crash when using flashinfer backend (#3207)
2025-02-02 15:36:07 -08:00
torchao_utils.py
Support loading of larger models with on-the-fly quantization (#3061)
2025-01-22 21:33:17 -08:00
vocab_parallel_embedding.py
feat: remove vllm distributed (#2907)
2025-01-17 22:31:51 +08:00
Powered by Gitea Version: 1.27.2 Page: 213ms Template: 3ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API