This website requires JavaScript.
Explore
Help
Sign In
chenchenghao
/
sglang
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
Actions
12
Packages
Projects
Releases
Wiki
Activity
Files
1a053a810c1957f2d32ac1e78d524a263d9186ac
sglang
/
docs
/
advanced_features
T
History
b8zhong
f374623fa9
[Refactor] Set
fp4-gemm-backend=auto
on SM100 and rename
fp4-gemm-backend
with
flashinfer_
prefix (
#17309
)
2026-01-19 20:09:07 +08:00
..
attention_backend.md
…
checkpoint_engine.md
…
cuda_graph_for_multi_modal_encoder.md
…
deterministic_inference.md
…
dp_for_multi_modal_encoder.md
…
epd_disaggregation.md
…
expert_parallelism.md
[Docs] minor update on ep docs (
#17242
)
2026-01-16 18:20:47 -08:00
forward_hooks.md
…
hicache_best_practices.md
…
hicache_design.md
…
hicache.rst
…
hyperparameter_tuning.md
…
lora.ipynb
[Feature] overlap LoRA weight loading with compute (
#15512
)
2026-01-19 10:43:17 +08:00
observability.md
…
pd_disaggregation.md
[NPU] upgrade npu mf_apater plugin (
#15853
)
2026-01-13 09:02:10 +08:00
pipeline_parallelism.md
…
quantization.md
[AMD][Quantization] Add
int4fp8_moe
online quantization on ROCm (
#7392
)
2026-01-14 01:44:40 -08:00
quantized_kv_cache.md
…
rfork.md
support non disturbing remote instance weight loader v2 (
#14997
)
2025-12-16 14:39:56 -08:00
separate_reasoning.ipynb
docs only add kimi k2 thinking and kimi linear (
#15789
)
2026-01-15 12:09:52 -05:00
server_arguments.md
[Refactor] Set
fp4-gemm-backend=auto
on SM100 and rename
fp4-gemm-backend
with
flashinfer_
prefix (
#17309
)
2026-01-19 20:09:07 +08:00
sgl_model_gateway.md
[model-gateway] Add Redis support as a history backend (
#16300
)
2026-01-11 01:03:00 -08:00
speculative_decoding.ipynb
…
structured_outputs_for_reasoning_models.ipynb
…
structured_outputs.ipynb
…
tool_parser.ipynb
…
vlm_query.ipynb
…