Let tail chunks co-batch with cache-hit requests (SGLANG_CP_PREFILL_MIX_CHUNKED)
专题 S1 (design: docs_internal/perf/prefill-compute-intensity-plan.md S1.0-S1.11). A batch containing a chunked prefill request has been forced to bs=1 by the CP gate, so every chunk of a long prompt monopolizes a forward while short-extend cache-hit continuations queue — the direct cause of the replay TTFT tail (p90 19.3s / p99 49.2s at 91.8% cache hit). Yet mixed chunk batches already occur today (a freshly-chunked request keeps earlier-admitted small requests), proving the CP forward path is mixed-chunk-safe; only admission was asymmetric. Three changes, the first flag-independent: - add_chunked_req now seeds the budget with the chunk's TRUE prefix (was 0), so the CP cached tally and the buffer estimator's mqa_logits k_rows see the chunk's footprint before any later request is admitted (landmine D1). - New SGLANG_CP_PREFILL_MIX_CHUNKED (default OFF): with it on, the gate admits requests after a chunked one and lets the existing CP caps (extend / cached / buffer, now correctly seeded) bound the batch — a FULL chunk still ends the scan by consuming the chunk-clamped extend cap; only a tail chunk leaves headroom. A chunked prefix that is not page-aligned (rare sub-page final-chunk tail) keeps its batch solo (the CP page-aligned split would fail-fast otherwise). - The symm staging capacity identity (admission extend cap + request slack == staging pages) is asserted when the flag is on, locking the coupling the design relies on (plan doc S1.4 I2). Tests: 4 new adder units (budget seeding; tail chunk admits followers; full chunk solo by budget; non-aligned prefix solo); the 8-rank byte-exactness scenario gains a chunk-shaped request (2048-token page-aligned carried prefix + 512 extend) — all four phases (legacy/v2/symm/prefetch) byte-identical on g0033. Known pre-existing cross-file pollution noted in problems.md P16. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -236,6 +236,12 @@ class Envs:
|
||||
# Total symm slab size override in MB (both halves). 0 = pool-derived
|
||||
# hard bound (logical pages x dense page unit, see design doc).
|
||||
SGLANG_CP_SHARED_KV_SYMM_HEAP_MB = EnvInt(0)
|
||||
# Allow a chunked prefill request (page-aligned carried prefix) to
|
||||
# co-batch with other requests under CP shared-KV bs>1: a tail chunk
|
||||
# leaves extend-cap headroom for short-extend cache-hit requests; a full
|
||||
# chunk still ends the scan by budget. See
|
||||
# docs_internal/perf/prefill-compute-intensity-plan.md S1.
|
||||
SGLANG_CP_PREFILL_MIX_CHUNKED = EnvBool(False)
|
||||
# NSA MQA logits are materialized as fp32 [q, k] buffers inside DeepGEMM.
|
||||
# Lower values split query rows more aggressively to cap peak temporary memory.
|
||||
SGLANG_NSA_MQA_LOGITS_FREE_MEM_FRACTION = EnvFloat(0.2)
|
||||
|
||||
Reference in New Issue
Block a user