This website requires JavaScript.
Explore
Help
Sign In
chenchenghao
/
sglang
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
Actions
12
Packages
Projects
Releases
Wiki
Activity
Files
6252ade98571c3374d7e7df3430a2bfbddfc5eb3
sglang
/
python
/
sglang
/
srt
/
layers
/
attention
T
History
HAI
6252ade985
revert BLOCK and num_warps on HIP (
#3722
)
2025-02-20 23:30:18 +08:00
..
triton_ops
revert BLOCK and num_warps on HIP (
#3722
)
2025-02-20 23:30:18 +08:00
__init__.py
Improve linear.py to load sharded weights & remove the dependency of Parameters from vllm (
#2784
)
2025-01-07 23:29:10 -08:00
double_sparsity_backend.py
Update Triton extend backend interface (
#3309
)
2025-02-05 18:12:22 +08:00
flashinfer_backend.py
Fix draft decode max batch size (
#3676
)
2025-02-18 23:03:26 +08:00
torch_native_backend.py
Support target model verification in the attention backend (
#2678
)
2024-12-30 22:58:55 -08:00
triton_backend.py
Fix draft decode max batch size (
#3676
)
2025-02-18 23:03:26 +08:00
vision.py
fix: apply cache size limit of attention mask for VisionAttention (
#3657
)
2025-02-19 20:16:48 +08:00