Files
sglang/docs/advanced_features/speculative_decoding.ipynb
T

24 KiB

Speculative Decoding

SGLang provides several speculative decoding options, including EAGLE-2/EAGLE-3, MTP, classic draft-model decoding, and an NGRAM-based variant. Our implementation aims to maximize speed and efficiency and is considered to be among the fastest in open-source LLM engines.

Summary

Jump to sections

Quick guidance

  • Best speed/quality (recommended): Use EAGLE-3 with --speculative-algorithm EAGLE3.
  • Strong default / broad compatibility: Use EAGLE-2 with --speculative-algorithm EAGLE.
  • Lower lm_head overhead for EAGLE-2: Enable FR-Spec with --speculative-token-map.
  • Model is MTP-enabled: Use MTP via speculative decoding (often with small speculative_num_steps/topk/num_draft_tokens, see the example section).
  • You have a smaller draft LLM: Use STANDALONE (--speculative-algorithm STANDALONE).
  • No extra model available: Use NGRAM (--speculative-algorithm NGRAM, CUDA-only).
  • Want overlap scheduler (experimental): Enable SpecV2 with SGLANG_ENABLE_SPEC_V2=True (requires --speculative-eagle-topk 1).

Method comparison (mini table)

Method Draft source Separate draft model? How to enable Notes / constraints
EAGLE-2 EAGLE draft model (feature drafting + tree) Typically yes --speculative-algorithm EAGLE + --speculative-draft-model-path ... Tune --speculative-num-steps, --speculative-eagle-topk, --speculative-num-draft-tokens
EAGLE-2 + torch.compile Same as EAGLE-2 Typically yes Add --enable-torch-compile (optionally --torch-compile-max-bs) Further kernel-level optimizations
EAGLE-2 + FR-Spec Same as EAGLE-2 + token subset Typically yes Add --speculative-token-map ... Reduces lm_head overhead with high-frequency token vocab
EAGLE-3 EAGLE3 draft model Yes --speculative-algorithm EAGLE3 + --speculative-draft-model-path ... Best throughput in the benchmark above
MTP Built-in multi-token heads (model-specific) Often no See Multi Token Prediction section Uses speculative workflow; draft path may be auto-handled for some models
STANDALONE Smaller draft LLM (token-level) Yes --speculative-algorithm STANDALONE + --speculative-draft-model-path ... Does not support --enable-dp-attention
SpecV2 (experimental) V2 workers + overlap scheduler N/A SGLANG_ENABLE_SPEC_V2=True Only supports --speculative-eagle-topk 1; applies to EAGLE, EAGLE3, STANDALONE
NGRAM Ngram cache from previous tokens No --speculative-algorithm NGRAM CUDA-only; no --enable-dp-attention; disables overlap scheduler & mixed chunked prefill

Performance Highlights

Please see below for the huge improvements on throughput for LLaMA-Instruct 3.1 8B tested on MT bench that can be achieved via EAGLE3 decoding. For further details please see the EAGLE3 paper.

Method Throughput (tokens/s)
SGLang (w/o speculative, 1x H100) 158.34 tokens/s
SGLang + EAGLE-2 (1x H100) 244.10 tokens/s
SGLang + EAGLE-3 (1x H100) 373.25 tokens/s

EAGLE Decoding

To enable EAGLE speculative decoding the following parameters are relevant:

  • speculative_draft_model_path: Draft model path/weights. Typically required for EAGLE/EAGLE3 and STANDALONE. For some MTP-enabled models, this can be omitted (SGLang may auto-handle/auto-fill it).
  • speculative_num_steps: Depth of autoregressive drafting. Increases speculation range but risks rejection cascades. Default is 5.
  • speculative_eagle_topk: Branching factor per step. Improves candidate diversity, will lead to higher acceptance rate, but more lead to higher memory/compute consumption. Default is 4.
  • speculative_num_draft_tokens: Maximum parallel verification capacity. Allows deeper tree evaluation but will lead to higher GPU memory usage. Default is 8.

These parameters are the same for EAGLE-2 and EAGLE-3.

You can find the best combinations of these parameters with bench_speculative.py.

In the documentation below, we set --cuda-graph-max-bs to be a small value for faster engine startup. For your own workloads, please tune the above parameters together with --cuda-graph-max-bs, --max-running-requests, --mem-fraction-static for the best performance.

EAGLE-2 decoding

You can enable EAGLE-2 decoding by setting --speculative-algorithm EAGLE and choosing an appropriate model.

In [ ]:
from sglang.test.doc_patch import launch_server_cmd
from sglang.utils import wait_for_server, print_highlight, terminate_process

import openai
In [ ]:
server_process, port = launch_server_cmd(
    """
python3 -m sglang.launch_server --model meta-llama/Llama-2-7b-chat-hf  --speculative-algorithm EAGLE \
    --speculative-draft-model-path lmsys/sglang-EAGLE-llama2-chat-7B --speculative-num-steps 3 \
    --speculative-eagle-topk 4 --speculative-num-draft-tokens 16 --cuda-graph-max-bs 8 --log-level warning
"""
)

wait_for_server(f"http://localhost:{port}")
In [ ]:
client = openai.Client(base_url=f"http://127.0.0.1:{port}/v1", api_key="None")

response = client.chat.completions.create(
    model="meta-llama/Llama-2-7b-chat-hf",
    messages=[
        {"role": "user", "content": "List 3 countries and their capitals."},
    ],
    temperature=0,
    max_tokens=64,
)

print_highlight(f"Response: {response}")
In [ ]:
terminate_process(server_process)

EAGLE-2 Decoding with torch.compile

You can also enable torch.compile for further optimizations and optionally set --torch-compile-max-bs:

In [ ]:
server_process, port = launch_server_cmd(
    """
python3 -m sglang.launch_server --model meta-llama/Llama-2-7b-chat-hf  --speculative-algorithm EAGLE \
    --speculative-draft-model-path lmsys/sglang-EAGLE-llama2-chat-7B --speculative-num-steps 5 \
        --speculative-eagle-topk 8 --speculative-num-draft-tokens 64 --mem-fraction 0.6 \
            --enable-torch-compile --torch-compile-max-bs 2 --log-level warning
"""
)

wait_for_server(f"http://localhost:{port}")
In [ ]:
client = openai.Client(base_url=f"http://127.0.0.1:{port}/v1", api_key="None")

response = client.chat.completions.create(
    model="meta-llama/Llama-2-7b-chat-hf",
    messages=[
        {"role": "user", "content": "List 3 countries and their capitals."},
    ],
    temperature=0,
    max_tokens=64,
)

print_highlight(f"Response: {response}")
In [ ]:
terminate_process(server_process)

EAGLE-2 Decoding via Frequency-Ranked Speculative Sampling

By employing a truncated high-frequency token vocabulary in the draft model, Eagle speculative decoding reduces lm_head computational overhead while accelerating the pipeline without quality degradation. For more details, checkout the paper.

In our implementation, set --speculative-token-map to enable the optimization. You can get the high-frequency token in FR-Spec from this model. Or you can obtain high-frequency token by directly downloading these token from this repo.

Thanks for the contribution from Weilin Zhao and Zhousx.

In [ ]:
server_process, port = launch_server_cmd(
    """
python3 -m sglang.launch_server --model meta-llama/Meta-Llama-3-8B-Instruct --speculative-algorithm EAGLE \
    --speculative-draft-model-path lmsys/sglang-EAGLE-LLaMA3-Instruct-8B --speculative-num-steps 5 \
    --speculative-eagle-topk 8 --speculative-num-draft-tokens 64 --speculative-token-map thunlp/LLaMA3-Instruct-8B-FR-Spec/freq_32768.pt \
    --mem-fraction 0.7 --cuda-graph-max-bs 2 --dtype float16  --log-level warning
"""
)

wait_for_server(f"http://localhost:{port}")
In [ ]:
client = openai.Client(base_url=f"http://127.0.0.1:{port}/v1", api_key="None")

response = client.chat.completions.create(
    model="meta-llama/Meta-Llama-3-8B-Instruct",
    messages=[
        {"role": "user", "content": "List 3 countries and their capitals."},
    ],
    temperature=0,
    max_tokens=64,
)

print_highlight(f"Response: {response}")
In [ ]:
terminate_process(server_process)

EAGLE-3 Decoding

You can enable EAGLE-3 decoding by setting --speculative-algorithm EAGLE3 and choosing an appropriate model.

In [ ]:
server_process, port = launch_server_cmd(
    """
python3 -m sglang.launch_server --model meta-llama/Llama-3.1-8B-Instruct  --speculative-algorithm EAGLE3 \
    --speculative-draft-model-path jamesliu1/sglang-EAGLE3-Llama-3.1-Instruct-8B --speculative-num-steps 5 \
        --speculative-eagle-topk 8 --speculative-num-draft-tokens 32 --mem-fraction 0.6 \
        --cuda-graph-max-bs 2 --dtype float16 --log-level warning
"""
)

wait_for_server(f"http://localhost:{port}")
In [ ]:
client = openai.Client(base_url=f"http://127.0.0.1:{port}/v1", api_key="None")

response = client.chat.completions.create(
    model="meta-llama/Meta-Llama-3.1-8B-Instruct",
    messages=[
        {"role": "user", "content": "List 3 countries and their capitals."},
    ],
    temperature=0,
    max_tokens=64,
)

print_highlight(f"Response: {response}")
In [ ]:
terminate_process(server_process)

Multi Token Prediction

We support MTP(Multi-Token Prediction) in SGLang by using speculative decoding. We use Xiaomi/MiMo-7B-RL model as example here (deepseek mtp usage refer to deepseek doc)

In [ ]:
server_process, port = launch_server_cmd(
    """
    python3 -m sglang.launch_server --model-path XiaomiMiMo/MiMo-7B-RL --host 0.0.0.0 --trust-remote-code \
    --speculative-algorithm EAGLE --speculative-num-steps 1 --speculative-eagle-topk 1 --speculative-num-draft-tokens 2 \
    --mem-fraction 0.5 --log-level warning
"""
)

wait_for_server(f"http://localhost:{port}")
In [ ]:
import requests

url = f"http://localhost:{port}/v1/chat/completions"

data = {
    "model": "XiaomiMiMo/MiMo-7B-RL",
    "messages": [{"role": "user", "content": "What is the capital of France?"}],
}

response = requests.post(url, json=data)
print_highlight(response.json())
In [ ]:
terminate_process(server_process)

Standalone Speculative Decoding (Small Draft Model)

Besides EAGLE/MTP, SGLang also supports token-level speculative decoding using a smaller draft model. Enable it with --speculative-algorithm STANDALONE and provide a draft model via --speculative-draft-model-path.

Relevant parameters:

  • --speculative-draft-model-path: Draft model weights (smaller than the target model).
  • --speculative-num-steps: Draft depth (how many steps the draft model runs autoregressively).
  • --speculative-eagle-topk: Branching factor (token candidates per step).
  • --speculative-num-draft-tokens: Verification capacity.

Note:

  • Standalone speculative decoding currently does not support --enable-dp-attention.
In [ ]:
server_process, port = launch_server_cmd(
    """
python3 -m sglang.launch_server --model Qwen/Qwen2.5-7B-Instruct --speculative-algorithm STANDALONE \
    --speculative-draft-model-path Qwen/Qwen2.5-1.5B-Instruct \
    --speculative-num-steps 4 --speculative-eagle-topk 2 --speculative-num-draft-tokens 7 \
    --cuda-graph-max-bs 8 --mem-fraction-static 0.7 --log-level warning
"""
)

wait_for_server(f"http://localhost:{port}")
In [ ]:
client = openai.Client(base_url=f"http://127.0.0.1:{port}/v1", api_key="None")

response = client.chat.completions.create(
    model="Qwen/Qwen2.5-7B-Instruct",
    messages=[
        {"role": "user", "content": "List 3 countries and their capitals."},
    ],
    temperature=0,
    max_tokens=64,
)

print_highlight(f"Response: {response}")
In [ ]:
terminate_process(server_process)

Speculative Decoding V2 (Overlap Scheduler)

SGLang provides an experimental Speculative Decoding V2 implementation that enables an overlap scheduler and uses V2 speculative workers (e.g. StandaloneWorkerV2, EAGLEWorkerV2).

To enable it, set the environment variable:

  • SGLANG_ENABLE_SPEC_V2=True

Notes:

  • SpecV2 currently only supports --speculative-eagle-topk 1. When SpecV2 is enabled, set --speculative-eagle-topk 1 explicitly.
  • If you explicitly set --speculative-eagle-topk > 1, the server will error. If you omit --speculative-eagle-topk, auto-tuning may pick topk > 1 for some models (e.g. Llama), which is not supported by SpecV2.
  • This applies to EAGLE, EAGLE3, and STANDALONE.
In [ ]:
server_process, port = launch_server_cmd(
    """
SGLANG_ENABLE_SPEC_V2=True python3 -m sglang.launch_server --model Qwen/Qwen2.5-7B-Instruct --speculative-algorithm STANDALONE \
    --speculative-draft-model-path Qwen/Qwen2.5-1.5B-Instruct \
    --speculative-num-steps 4 --speculative-eagle-topk 1 --speculative-num-draft-tokens 5 \
    --cuda-graph-max-bs 8 --mem-fraction-static 0.7 --log-level warning
"""
)

wait_for_server(f"http://localhost:{port}")
In [ ]:
client = openai.Client(base_url=f"http://127.0.0.1:{port}/v1", api_key="None")

response = client.chat.completions.create(
    model="Qwen/Qwen2.5-7B-Instruct",
    messages=[
        {"role": "user", "content": "List 3 countries and their capitals."},
    ],
    temperature=0,
    max_tokens=64,
)

print_highlight(f"Response: {response}")
In [ ]:
terminate_process(server_process)

Ngram Speculative Decoding

SGLang also supports ngram-based speculative decoding (no separate draft model). It retrieves draft tokens from an ngram cache built from previously generated tokens, and then verifies them with the target model.

Enable it with:

  • --speculative-algorithm NGRAM

Common parameters:

  • --speculative-num-draft-tokens: Number of draft tokens verified per step.
  • --speculative-ngram-min-match-window-size / --speculative-ngram-max-match-window-size: Matching window range.
  • --speculative-ngram-min-bfs-breadth / --speculative-ngram-max-bfs-breadth: BFS breadth range.
  • --speculative-ngram-branch-length: How many recent tokens to insert into the cache.
  • --speculative-ngram-capacity: Cache capacity.

Notes:

  • Ngram speculative decoding only supports CUDA.
  • It currently does not support --enable-dp-attention.
  • It disables the overlap scheduler and mixed chunked prefill.
  • Optional: set SGLANG_NGRAM_FORCE_GREEDY_VERIFY=True to force greedy verification.
In [ ]:
server_process, port = launch_server_cmd(
    """
python3 -m sglang.launch_server --model Qwen/Qwen2.5-7B-Instruct --speculative-algorithm NGRAM \
    --speculative-num-draft-tokens 16 \
    --speculative-ngram-max-match-window-size 12 --speculative-ngram-max-bfs-breadth 10 \
    --cuda-graph-max-bs 8 --mem-fraction-static 0.8 --log-level warning
"""
)

wait_for_server(f"http://localhost:{port}")
In [ ]:
client = openai.Client(base_url=f"http://127.0.0.1:{port}/v1", api_key="None")

response = client.chat.completions.create(
    model="Qwen/Qwen2.5-7B-Instruct",
    messages=[
        {"role": "user", "content": "List 3 countries and their capitals."},
    ],
    temperature=0,
    max_tokens=64,
)

print_highlight(f"Response: {response}")
In [ ]:
terminate_process(server_process)

References

EAGLE process is as follows:

  • Within EAGLE the draft model predicts the next feature vector, i.e. the last hidden state of the original LLM, using the feature sequence (f_1, ..., f_k) and the token sequence (t_2, ..., t_{k+1}).
  • The next token is then sampled from p_{k+2}=\text{LMHead}(f_{k+1}). Afterwards, the two sequences are extended in a tree style—branching out multiple potential continuations, with the branching factor per step controlled by the speculative_eagle_topk parameter—to ensure a more coherent connection of context, and are given as input again.
  • EAGLE-2 additionally uses the draft model to evaluate how probable certain branches in the draft tree are, dynamically stopping the expansion of unlikely branches. After the expansion phase, reranking is employed to select only the top speculative_num_draft_tokens final nodes as draft tokens.
  • EAGLE-3 removes the feature prediction objective, incorporates low and mid-layer features, and is trained in an on-policy manner.

This enhances drafting accuracy by operating on the features instead of tokens for more regular inputs and passing the tokens from the next timestep additionally to minimize randomness effects from sampling. Furthermore the dynamic adjustment of the draft tree and selection of reranked final nodes increases acceptance rate of draft tokens further. For more details see EAGLE-2 and EAGLE-3 paper.

For guidance how to train your own EAGLE model please see the EAGLE repo.