This website requires JavaScript.
Explore
Help
Sign In
chenchenghao
/
sglang
Watch
1
Star
0
Fork
0
You've already forked sglang
Code
Issues
Pull Requests
Actions
11
Packages
Projects
Releases
Wiki
Activity
Files
732fc8e405df5f4891f5f61621df3e0d7bf8b1e5
sglang
/
python
/
sglang
/
srt
/
layers
/
moe
History
Peng Zhang
191d836ff6
fix: minor fix for modelopt weight load compatibility (
#7953
)
2025-07-11 14:20:58 -07:00
..
ep_moe
[feature]Ascend quantization support (
#7791
)
2025-07-10 09:17:37 -07:00
fused_moe_triton
fix: minor fix for modelopt weight load compatibility (
#7953
)
2025-07-11 14:20:58 -07:00
cutlass_moe_params.py
[CUTLASS-FP4-MOE] Introduce CutlassMoEParams class for easy initialization of Cutlass Grouped Gems Metadata (
#6887
)
2025-06-05 13:13:14 -07:00
cutlass_moe.py
Add a CUDA kernel for fusing mapping and weighted sum for MoE. (
#6916
)
2025-06-07 15:24:39 -07:00
cutlass_w4a8_moe.py
feat: support DeepSeek-R1-W4AFP8 model with ep-moe mode (
#7762
)
2025-07-07 14:47:21 -07:00
fused_moe_native.py
[CPU] [BF16] Call fused_experts_cpu, weight_packed_linear and bmm_cpu kernel in DeepSeek model (
#6641
)
2025-06-25 01:43:33 -07:00
router.py
Improve streaming, log_level, memory report, weight loading, and benchmark script (
#7632
)
2025-06-29 23:16:19 -07:00
topk.py
[feature]Ascend quantization support (
#7791
)
2025-07-10 09:17:37 -07:00