This website requires JavaScript.
Explore
Help
Sign In
chenchenghao
/
sglang
Watch
1
Star
0
Fork
0
You've already forked sglang
Code
Issues
Pull Requests
Actions
12
Packages
Projects
Releases
Wiki
Activity
Files
6d4030890568aaa7b73dbfc828b75787504f79a0
sglang
/
python
/
sglang
/
srt
/
speculative
History
Stefan He
6c18ab46a2
[Qwen3-Next] switch to triton and cache conv states to accelerate MTP from 300 tok/s to 341 tok/s (
#10335
)
...
Co-authored-by: Binyao Jiang <
byjiang1996@gmail.com
>
2025-09-11 11:59:48 -07:00
..
build_eagle_tree.py
Add treemask mode to build_eagle_tree & release sgl-kernel 0.2.3 (
#7756
)
2025-07-05 12:17:05 -07:00
eagle_draft_cuda_graph_runner.py
Qwen2.5-VL eagle3 infer (
#8801
)
2025-09-07 20:44:34 -07:00
eagle_draft_extend_cuda_graph_runner.py
Standalone speculative decoding (
#10090
)
2025-09-07 20:55:09 -07:00
eagle_utils.py
[MTP] Force greedy sampling on AMD (
#9127
)
2025-08-22 11:14:43 -07:00
eagle_worker.py
[Qwen3-Next] switch to triton and cache conv states to accelerate MTP from 300 tok/s to 341 tok/s (
#10335
)
2025-09-11 11:59:48 -07:00
spec_info.py
Standalone speculative decoding (
#10090
)
2025-09-07 20:55:09 -07:00
standalone_worker.py
Standalone speculative decoding (
#10090
)
2025-09-07 20:55:09 -07:00