Logo
Explore Help
Sign In
chenchenghao/sglang
1
0
Fork 0
You've already forked sglang
Code Issues Pull Requests Actions 11 Packages Projects Releases Wiki Activity
Files
9a00e6f453e764c0b286e2a62f652a1202c0bf9c
sglang/python/sglang/srt/layers
History
HAI f35cb46cc3 ROCm: Fix MoE padding for none FP8 cases (#2111)
2024-11-21 12:23:21 -08:00
..
attention
Enable overlap scheduler by default for the triton attention backend (#2105)
2024-11-20 02:58:35 -08:00
fused_moe
ROCm: Fix MoE padding for none FP8 cases (#2111)
2024-11-21 12:23:21 -08:00
quantization
fix black in pre-commit (#1940)
2024-11-08 07:42:47 +08:00
activation.py
feat: update torch 2.5.1 (#2069)
2024-11-18 21:29:13 +08:00
custom_op_util.py
feat: update torch 2.5.1 (#2069)
2024-11-18 21:29:13 +08:00
layernorm.py
feat: update torch 2.5.1 (#2069)
2024-11-18 21:29:13 +08:00
linear.py
Update vllm to 0.6.3 (#1711) (#1720)
2024-10-19 20:45:41 -07:00
logits_processor.py
Fix illegal memory access in overlap mode & Use more fused triton kernels for building meta data (#2051)
2024-11-16 16:14:23 -08:00
pooler.py
Rename InputMetadata -> ForwardBatch (#1543)
2024-09-30 02:41:11 -07:00
radix_attention.py
Simplify flashinfer dispatch (#1552)
2024-10-01 00:28:42 -07:00
rotary_embedding.py
Qwen2vl support cuda graph and disable radix cache (#1780)
2024-10-25 10:45:17 -04:00
sampler.py
Make constrained decoding work for overlap scheduler (#2095)
2024-11-19 15:04:43 -08:00
torchao_utils.py
Error out when torchao-config option is not recognized (#2107)
2024-11-20 17:37:28 -08:00
vocab_parallel_embedding.py
fix black in pre-commit (#1940)
2024-11-08 07:42:47 +08:00
Powered by Gitea Version: 1.25.5 Page: 80ms Template: 10ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API