Group cache-hit prefills into dense batches (SGLANG_CP_PREFILL_AFFINITY_GROUP)
专题 S4 (design: docs_internal/perf/prefill-compute-intensity-plan.md S4, amended). Under FCFS a cold request joining a warm-led batch turns a 1-2s cache-hit forward into a 5-10s one, splitting the warm work into the 新-cache-新 pattern. The policy prevents exactly that one thing: - WARM candidates always admit (into a cold-led batch they are free density — the cold extend dominates the forward anyway). - COLD admits into an empty or cold-led batch (small colds co-batch today; the FCFS head always starts a batch so the queue keeps moving). - COLD into a WARM-led batch is skipped, bounded by a per-pass window (W=16 skips), a head defer count (K=3 passes) and an age bound (T=5s). On any bound the scan STOPS instead of force-admitting: the cold waits for the same forward either way, but leads its own clean batch next pass instead of polluting this one. The skip is strictly post-match / pre-admit (after init_next_round_input, before add_one_req): no lock, no allocation, no budget mutation to unwind, and re-matching a skipped candidate next pass is exactly what the scan already does after a cap rejection. Classification is the in-scan match result (device prefix + host hit vs a 64-token floor) — under FCFS+L2 no pre-scan signal exists, so this adds zero matching work for inspected candidates. Disabled wholesale under priority scheduling (the skip must not reorder across priority classes). Three amendments vs the design draft, reasoned in the decision-table docstring: cold+cold-led admits (STOP would regress today's small-cold co-batching); starved heads STOP rather than force-admit (clean batch boundaries at identical latency); priority interaction handled by disabling rather than per-request comparison. Decision logic is a pure function with table + bounds unit tests (28/28 adder suite green). Default OFF. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -242,6 +242,23 @@ class Envs:
|
||||
# chunk still ends the scan by budget. See
|
||||
# docs_internal/perf/prefill-compute-intensity-plan.md S1.
|
||||
SGLANG_CP_PREFILL_MIX_CHUNKED = EnvBool(False)
|
||||
# Cache-affinity batch formation (plan doc S4): keep FCFS, but skip a
|
||||
# COLD candidate while the batch is warm-led so cache-hit requests group
|
||||
# into one dense forward instead of being split by a cold 60K chunk.
|
||||
# Bounded by the window/defer/age knobs below; disabled automatically
|
||||
# under priority scheduling.
|
||||
SGLANG_CP_PREFILL_AFFINITY_GROUP = EnvBool(False)
|
||||
# Max cold candidates skipped per batch-formation pass (scan depth bound;
|
||||
# each skip costs one re-match next pass).
|
||||
SGLANG_CP_PREFILL_AFFINITY_WINDOW = EnvInt(16)
|
||||
# Max consecutive passes the FCFS head may be deferred before the scan
|
||||
# stops grouping and lets it lead the next (empty) batch.
|
||||
SGLANG_CP_PREFILL_AFFINITY_MAX_DEFER = EnvInt(3)
|
||||
# Absolute age bound (seconds) with the same effect as MAX_DEFER; must be
|
||||
# well under SGLANG_REQ_WAITING_TIMEOUT when that is set.
|
||||
SGLANG_CP_PREFILL_AFFINITY_MAX_AGE_S = EnvFloat(5.0)
|
||||
# Warm/cold threshold in tokens on (device prefix hit + host hit).
|
||||
SGLANG_CP_PREFILL_AFFINITY_WARM_FLOOR = EnvInt(64)
|
||||
# NSA MQA logits are materialized as fp32 [q, k] buffers inside DeepGEMM.
|
||||
# Lower values split query rows more aggressively to cap peak temporary memory.
|
||||
SGLANG_NSA_MQA_LOGITS_FREE_MEM_FRACTION = EnvFloat(0.2)
|
||||
|
||||
@@ -643,6 +643,9 @@ class Req(ReqDllmMixin):
|
||||
self.last_host_node: Any = None
|
||||
self.last_host_backup_node: Any = None
|
||||
self.host_hit_length = 0
|
||||
# Cache-affinity scheduling (plan doc S4): consecutive passes this
|
||||
# request was deferred as the FCFS head while a warm batch formed.
|
||||
self.affinity_defer_count = 0
|
||||
self.prefix_match_deferred_by_pending_backup = False
|
||||
self.cp_hicache_prepared_backup = None
|
||||
# Tokens loaded from storage backend (L3) during prefetch for this request
|
||||
|
||||
@@ -101,6 +101,60 @@ def _cp_shared_kv_bs_gt1_admission_debug(
|
||||
logger.info("[CP_SHARED_KV_BS_GT1_DEBUG] event=%s " + message, key, *args)
|
||||
|
||||
|
||||
class AffinityDecision(Enum):
|
||||
"""Outcome of the cache-affinity scan decision (plan doc S4)."""
|
||||
|
||||
ADMIT = auto()
|
||||
SKIP_COLD = auto()
|
||||
STOP = auto()
|
||||
|
||||
|
||||
def decide_cp_prefill_affinity(
|
||||
*,
|
||||
is_warm: bool,
|
||||
batch_empty_for_affinity: bool,
|
||||
batch_warm_led: bool,
|
||||
is_head: bool,
|
||||
head_defer_count: int,
|
||||
head_age_s: float,
|
||||
window_used: int,
|
||||
) -> AffinityDecision:
|
||||
"""Cache-affinity batch-formation decision (pure; plan doc S4, amended).
|
||||
|
||||
The policy prevents exactly one thing: a COLD candidate joining a
|
||||
WARM-led batch (one cold 60K extend turns a 1-2s warm forward into a
|
||||
5-10s one). Everything else admits:
|
||||
|
||||
- WARM always admits — into a warm-led batch it grows the dense group;
|
||||
into a cold-led batch it is free density (the cold extend dominates
|
||||
the forward time anyway).
|
||||
- COLD into an empty batch admits — the FCFS head always starts a
|
||||
batch, so the queue keeps moving and a deferred cold leads the very
|
||||
next batch.
|
||||
- COLD into a COLD-led batch admits — small cold requests co-batch
|
||||
today (the CP caps bound them); refusing would be a regression.
|
||||
- COLD into a WARM-led batch is skipped, bounded three ways: the
|
||||
per-pass window, and for the FCFS head a defer count and an age
|
||||
bound. On any bound tripping the decision is STOP (end the warm
|
||||
batch cleanly) rather than admitting the cold into it: the cold
|
||||
waits for the same forward either way, but leads its own clean batch
|
||||
next pass instead of polluting this one.
|
||||
"""
|
||||
|
||||
if is_warm:
|
||||
return AffinityDecision.ADMIT
|
||||
if batch_empty_for_affinity or not batch_warm_led:
|
||||
return AffinityDecision.ADMIT
|
||||
if window_used >= int(envs.SGLANG_CP_PREFILL_AFFINITY_WINDOW.get()):
|
||||
return AffinityDecision.STOP
|
||||
if is_head and (
|
||||
head_defer_count >= int(envs.SGLANG_CP_PREFILL_AFFINITY_MAX_DEFER.get())
|
||||
or head_age_s >= float(envs.SGLANG_CP_PREFILL_AFFINITY_MAX_AGE_S.get())
|
||||
):
|
||||
return AffinityDecision.STOP
|
||||
return AffinityDecision.SKIP_COLD
|
||||
|
||||
|
||||
class CacheAwarePolicy(Enum):
|
||||
"""Scheduling policies that are aware of the tree cache."""
|
||||
|
||||
|
||||
@@ -69,6 +69,7 @@ from sglang.srt.eplb.expert_distribution import get_global_expert_distribution_r
|
||||
from sglang.srt.layers.attention.mamba.ops import (
|
||||
initialize_mamba_selective_state_update_backend,
|
||||
)
|
||||
from sglang.srt.layers.attention.nsa.utils import is_nsa_prefill_cp_in_seq_split
|
||||
from sglang.srt.layers.dp_attention import (
|
||||
compute_dp_attention_world_info,
|
||||
get_attention_cp_group,
|
||||
@@ -163,8 +164,10 @@ from sglang.srt.managers.scheduler_graceful_shutdown_mixin import (
|
||||
)
|
||||
from sglang.srt.managers.schedule_policy import (
|
||||
AddReqResult,
|
||||
AffinityDecision,
|
||||
PrefillAdder,
|
||||
SchedulePolicy,
|
||||
decide_cp_prefill_affinity,
|
||||
)
|
||||
from sglang.srt.managers.scheduler_dp_attn_mixin import SchedulerDPAttnMixin
|
||||
from sglang.srt.managers.scheduler_input_blocker import SchedulerInputBlocker
|
||||
@@ -2562,6 +2565,20 @@ class Scheduler(
|
||||
running_loras = {req.lora_id for req in self.running_batch.reqs}
|
||||
|
||||
# Get requests from the waiting queue to a new prefill batch
|
||||
# Cache-affinity grouping (plan doc S4): active only for the CP
|
||||
# shared-KV bs>1 prefill context, and never under priority
|
||||
# scheduling (the skip must not reorder across priority classes —
|
||||
# disabled wholesale as the conservative guard).
|
||||
affinity_on = (
|
||||
envs.SGLANG_CP_PREFILL_AFFINITY_GROUP.get()
|
||||
and not self.enable_priority_scheduling
|
||||
and self.server_args.enable_cp_shared_kv_prefill_bs_gt1
|
||||
and is_nsa_prefill_cp_in_seq_split()
|
||||
)
|
||||
affinity_window_used = 0
|
||||
affinity_head_pending = True # first classified candidate = FCFS head
|
||||
affinity_batch_warm_led: Optional[bool] = None
|
||||
|
||||
for req in self.waiting_queue:
|
||||
if self.enable_lora and req.lora_id not in running_loras:
|
||||
if self.enable_lora_overlap_loading:
|
||||
@@ -2606,11 +2623,54 @@ class Scheduler(
|
||||
)
|
||||
|
||||
req.init_next_round_input(self.tree_cache)
|
||||
|
||||
if affinity_on:
|
||||
# Decision strictly post-match / pre-admit: a skipped
|
||||
# candidate has no lock, no allocation, no budget mutation
|
||||
# to unwind, and re-matching it next pass is exactly what
|
||||
# the scan does today after a cap rejection.
|
||||
req_is_warm = (
|
||||
len(req.prefix_indices)
|
||||
+ int(getattr(req, "host_hit_length", 0) or 0)
|
||||
) >= int(envs.SGLANG_CP_PREFILL_AFFINITY_WARM_FLOOR.get())
|
||||
is_head = affinity_head_pending
|
||||
affinity_head_pending = False
|
||||
decision = decide_cp_prefill_affinity(
|
||||
is_warm=req_is_warm,
|
||||
batch_empty_for_affinity=(
|
||||
len(adder.can_run_list)
|
||||
- (1 if adder._chunked_req_in_batch is not None else 0)
|
||||
== 0
|
||||
),
|
||||
batch_warm_led=bool(affinity_batch_warm_led),
|
||||
is_head=is_head,
|
||||
head_defer_count=req.affinity_defer_count,
|
||||
head_age_s=(
|
||||
time.perf_counter()
|
||||
- req.time_stats.wait_queue_entry_time
|
||||
if req.time_stats.wait_queue_entry_time > 0.0
|
||||
else 0.0
|
||||
),
|
||||
window_used=affinity_window_used,
|
||||
)
|
||||
if decision is AffinityDecision.SKIP_COLD:
|
||||
if is_head:
|
||||
req.affinity_defer_count += 1
|
||||
affinity_window_used += 1
|
||||
continue
|
||||
if decision is AffinityDecision.STOP:
|
||||
break
|
||||
|
||||
num_admitted_before = len(adder.can_run_list)
|
||||
res = adder.add_one_req(
|
||||
req,
|
||||
has_chunked_req=(self.chunked_req is not None),
|
||||
truncation_align_size=self.truncation_align_size,
|
||||
)
|
||||
if affinity_on and len(adder.can_run_list) > num_admitted_before:
|
||||
req.affinity_defer_count = 0
|
||||
if affinity_batch_warm_led is None:
|
||||
affinity_batch_warm_led = req_is_warm
|
||||
|
||||
if self.enable_lora:
|
||||
running_loras.add(req.lora_id)
|
||||
|
||||
Reference in New Issue
Block a user