93fe86477969e2f039e238744bc7c1055ffa92aa
2298
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
07c2b9bc57 |
2b.2: B1 collective-free shared-pool evict-to-fit + delete the #5 deadlock
Replaces the 2b.1b shared-L2 capacity-miss SKIP with a deterministic collective-free evict-to-fit, and deletes the #5 deadlock. This is the gate that makes --enable-cp-shared-physical-l2-hicache safe in a real run. evict-to-fit (_reserve_write_cp_shared_l2_evict_to_fit, hiradix): on a shared-L2 reserve capacity miss -- which raises identically on every CP rank over the replicated free list -- snapshot the plannable shared-L2 host leaves (excluding the node being backed up) in the replicated SLRU order (_cp_host_evict_key = Phase-1 logical clock + node.id total-order tiebreak), release the coldest one at a time (each a replicated free-list mutation via the proven _evict_cp_host_for_write_admission teardown -> evict_cp_host -> allocator.release) and retry the pooled reserve after each, stopping at the first fit. Fragmentation-correct (free coalesces in _return_range). Identical victim set + order + deterministic reserve => every rank evicts the same victims and stops at the same point (design Thm 1), so NO collective is needed. The snapshot is one-level by design (a parent promoted to a leaf mid-loop is not re-pushed; deep cascades skip-and-retry next tick -- missed identically on all ranks). Exhaustion -> loud rate-limited skip (transient: pinned/in-flight objects hold the pool), never a hang. #5 deadlock DELETED: stripped the synchronize_across_ranks all_reduce machinery (the `while len(heap) and not all_ranks_done()` ReduceOp.MIN host_evict_done_min closure + the param) from _evict_host_for_physical_slots -- never reached (no caller passed True) and the headline CP-deadlock shape. evict_host(num_tokens) still works (single-arg, equivalent non-sync loop). Opus adversarial review = SHIP (all 7 findings PASS: determinism [_cp_host_evict_key purely replicated, pin_expiry dormant under CP], teardown byte-equivalence + no double-free, no infinite recursion + clean abort-on-partial, snapshot safety [no ancestor/descendant among host leaves], bounded termination, clean #5 strip, flag-off path untouched). Tests: new test_eight_ranks_evict_to_fit_stays_identical (fill->miss->release-in-order->retry -> identical placement, exercises coalescing) + 88/88 pool suite (syh-dev-new) + import smoke. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0d87bc576f |
2b dead-island GC: collapse pooled-L2 allocator to the B1 commit model
The 2b.1b adopt-then-strip left the entire option-A commit-quorum apparatus
orphaned: B1's writing_check ReduceOp.MIN frontier (mark_object_committed) is
the all-ranks-done consensus, so the per-(rank,layer,payload) gather quorum it
replaced is information-redundant (design sec 2.4 Corollary). This GC removes
the apparatus and collapses the CpSharedL2PageAllocator commit model to exactly
{_ranges_by_object, _committed_objects, _free_by_payload} -- precisely what
placement_digest() hashes -- so the cross-rank digest assert now covers the
whole model and the dead quorum is structurally unrepresentable.
Removed (all proven dead under B1, confirmed by a full tree caller-sweep +
opus adversarial review = SHIP):
- allocator: commit_layer, _has_full_commit, _expected_layers_for_object_payload,
adopt_reserved_range, set_object_required_payloads; fields _commits_by_object,
_expected_ranks, _adopted_ranges, _required_payloads_by_object,
_expected_layers_by_object, _required_payloads, _expected_layers; ctor args
expected_ranks/expected_layers/required_payloads; module fns
broadcast_cp_shared_l2_decision, gather_cp_shared_l2_{preflight,commits,commit}.
reserve/split_committed_object/_drop_object simplified to ranges+committed-bit.
- cache_controller: the 3 dead _commit_cp_shared_l2_* fns, 3 imports, all 5
option-A ctor params (cp_shared_l2_{cpu_group,source_rank,broadcast_fn,
preflight_fn,commit_fn}), the per-object commit-contract builder/setter/call.
- hiradix: allocator construction updated (3 args dropped); live _split_node
caller and _get_cp_shared_l2_rank_and_group@167 intact.
- tests: split tests migrated to mark_object_committed; 4 dead-symbol tests removed.
Lost __init__ payload-has-slab validation is covered by reserve()'s own
ValueError (fail-loud at first write). Default (flag-off) path unchanged: the
allocator is never constructed when enable_cp_shared_physical_l2_hicache=False.
Net -580 LOC. Validated: 87/87 pool suite (syh-dev-new, torch) + 31/31 core
classes locally + edited-module import smoke. Out of scope (flagged separately):
the CpSharedL2NodeMetadata.{required_payloads,committed_payload_layers} vestige.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
01029c8df4 |
2b.1b: B1 collective-free pooled-L2 write-through (default-off)
The B1 write-through: every CP rank runs the IDENTICAL deterministic CpSharedL2PageAllocator.reserve over the replicated event stream -> identical placement with NO broadcast/gather; commit rides the writing_check ReduceOp.MIN frontier (mark_object_committed) instead of a per-(rank,layer,payload) gather quorum. cache_controller.py: _reserve_write_cp_shared_l2 stripped to every-rank reserve (dropped preflight#2/rank0-gate/broadcast#3/adopt/missing-payloads; per-object contract kept MF3; idempotent abort-on-partial SF3); _submit_write_cp_layer_states drops the 3 gather-commit calls (per-layer D2H + all-payload done-gate MF1 + per-node ack kept); submit_write_cp_all_layer drops both fallback gathers (zero-owned keeps its ack MF2). hiradix_cache.py: 3-way merge of l2_pooling shared-L2 init (clean, 0-conflict; Phase-1 clock + E1 + reserve-budget preserved) builds the allocator/slab when the flag is on and passes it to HiCacheController WITHOUT collective fns. The #5 deadlock removed: shared_l2_capacity reserve-miss now SKIPS the backup (reactive, rank-uniform) instead of _evict_host_for_physical_slots(synchronize_across_ranks=True) -- the deterministic shared-pool evict is 2b.2. mark_object_committed wired at _commit_pending_backup (the MIN commit). New _cp_assert_placement_replicated (SGLANG_CP_HICACHE_PLACEMENT_ASSERT) in check_hicache_events: MIN/MAX-reduce placement_digest across the CP group, fail-loud on divergence (Theorem 1 runtime gate; rank-uniform entry, cannot deadlock). environ.py: SGLANG_CP_HICACHE_PLACEMENT_ASSERT (EnvBool, default off). Tests: 8-rank reserve-determinism (identical placement across 8 ranks + divergence-detectable) + full pool suite 91/91 + import smoke on g0033 syh-dev-new. TWO opus adversarial reviews both SHIP (cache_controller strip + hiradix wiring). NOTES: default-off (--enable-cp-shared-physical-l2-hicache); validated for write_through policy (prod default). Capacity-miss currently SKIPS (B1 shared-pool evict = 2b.2) so the flag must stay off in real runs until 2b.2. Dead gather island (3 fns + 3 imports, opus-confirmed inert) pending GC follow-up. Design+proof: docs_internal/cp_hicache_2b_pooled_l2_b1_design.md sec 2.4/9. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
50bc923ad6 |
2b.1b step 1: mark_object_committed (B1 MIN-driven commit)
Add CpSharedL2PageAllocator.mark_object_committed(object_key) -- the B1 commit path that marks an object committed directly, bypassing the per-(rank,layer,payload) commit_layer/_has_full_commit quorum. Under B1 the all-ranks-done consensus is the writing_check ReduceOp.MIN frontier (the same barrier that releases the node write lock), so every rank calls this with the same object_key at the MIN commit point and the committed set transitions rank-uniformly (feeds placement_digest). Resolves opus-review MF4: split_committed_object stays coherent for a B1-committed object (no per-layer commit facts) -- it adds children to _committed_objects and the empty _commits_by_object is harmless because B1 never calls _has_full_commit. 7 unit tests (pure-Python): commit-without-quorum, idempotent, unknown-raises, stat bump, digest reflects committed set + lockstep equality, split-of-marked-object coherent, release-after-mark frees pages. Full suite 88/88 green on g0033. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7850aab1a2 |
2b.1a: placement_digest gate on CpSharedL2PageAllocator (B1 proof obligation)
Add CpSharedL2PageAllocator.placement_digest() -- a deterministic, order-independent SHA-256 of the full replicated allocator state (free list + per-object ranges + committed set). This is the B1 proof-obligation gate from the 2b design (sec 2.4, Theorem 1): the write-through (2b.1) cross-rank assert hashes this each tick and EQ-checks it across CP ranks, turning placement-determinism into a fail-loud runtime invariant -- any reserve/release fired on a subset of ranks or with a per-rank arg changes the digest. 6 unit tests (pure-Python, no torch): determinism, changes-on-reserve, identical-op-sequence => identical-digest, divergent-op => different-digest, reserve+release restores digest (free-list coalescing), placement-order sensitivity. Full suite 81/81 green on g0033 syh-dev-new. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7de882b622 |
2b.0: adopt l2_pooling pooled-L2 mechanics (default-off)
Adopt the ownerless shared-physical-L2 mechanics from origin/l2_pooling onto the Phase-1 branch, behind --enable-cp-shared-physical-l2-hicache (default off -> behavior-neutral). Foundation for the B1 collective-free allocator (2b.1/2b.2) and the L3 durable floor (Phase 3); no coordination yet. - cp_shared_l2_pool.py (new): CpSharedL2PageAllocator (deterministic first-fit, capacity charged once, no x cp_size), hugetlbfs-2M slab primitives, position-indexed object ranges, commit-quorum data model. - memory_pool_host.py: SharedHostTensor* slab orchestration (3-way merged; fix-output's net change to this file is blank-line-only -> additions-only, no content lost). - server_args.py: flags + _handle_cp_shared_physical_l2_hicache_validation (3-way merged clean with Phase-1 assert + fix-output EnvField). - test_cp_shared_l2_pool.py: 75 unit tests (allocator/slab/quorum), green on g0033 syh-dev-new; envs stub provides SGLANG_REQ_WAITING_TIMEOUT (read by Phase-1 validation). Design + B1 proof: docs_internal/cp_hicache_2b_pooled_l2_b1_design.md (sec 2.4). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8f1e85a992 |
CP HiCache: replicated-clock SLRU eviction, collective-safe (Phase 1)
Make CP shared-KV HiCache eviction collective-free and scan-resistant by removing every per-rank wall-clock input from the eviction/admission decisions (any such input desyncs the must-be-replicated victim set / batch across the 8 CP ranks and can deadlock the collective-coupled writeback/load-back). - Replicated logical clock: TreeNode.next_access_time() (a process-global monotonic counter, mirrors mamba/swa's get_last_access_time) replaces time.monotonic() as the source of last_access_time at every radix bump site (radix_cache __init__/match/insert; hiradix CP match/insert/_insert_helper_host; reset() zeroes it). creation_time/pin_expiry intentionally stay wall-clock. Because match/insert events are replicated (reqs broadcast from rank 0 over the replicated tree), last_access_time is now identical on every rank. Unique ints also give a strict total order (no LRU tie ambiguity). No duration arithmetic reads last_access_time, so non-CP LRU is unchanged; mamba/swa use their own TreeNode and are untouched. - CpReplicatedSLRUStrategy = (is_protected[hit>=2], last_access_time, node.id): scan-resistant (cold one-shots stay probationary, evicted before reused/ protected prefixes), recency-within-segment ages out cold (not LFU), node.id is the deterministic total-order tiebreak (heaps never compare TreeNodes). Overridden for CP so it reaches every get_priority eviction site; the host write-admission key _cp_host_evict_key is switched from FIFO-by-id to it. - Co-fixes for the same per-rank-wall-clock class: pin_prefix is forbidden under CP (its pin_expiry victim-eligibility check is per-rank wall-clock; SLRU scan-resistance covers the benefit); the affinity head_age_s admission input is disabled (=0.0, relying on the replicated head_defer_count bound); and a server_args fail-fast guards CP shared-KV against SGLANG_REQ_WAITING_TIMEOUT>0 (its waiting-queue pruning is per-rank wall-clock). Tests: test_evict_policy.py adds TestCpReplicatedSLRUStrategy (scan-resistance, recency, total-order, full ordering) -- 30 passed in the dev-cu13 container; test_radix_cache_unit.py adds logical-clock determinism/total-order tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8d73919d02 |
Gate CP HiCache write-backup on hard deficit, not the free-room watermark
A long cold-request workload flooded the prefill log (~177MB / 32K lines in 8min, all 8 CP ranks) with cp_host_reservation_plan_insufficient + prepare_write_backup_reservation_failed even though the host pool was <42% used and the writes physically fit. Root cause: _free_room_deficit folds the proactive 20% free-room target (hicache_host_free_room_ratio) into the admission deficit once a lane dips below the 10% trigger. _evict_cp_host_for_write_admission then failed -- and evicted nothing -- whenever it could not reclaim that full 20% target, which is impossible while the residual is locked/pending. The lane never drained, so every subsequent backup re-issued the same impossible demand and re-logged on every tick. The reservation failure is itself a graceful skip, so there is no crash; the symptom is the log storm. Gate write-backup admission on the HARD deficit -- max(0, required-available) over the target and draft lanes (draft_available is 2**62 with no draft pool, a no-op) -- and evict toward the watermark best-effort, admitting whenever the write itself fits. This drains the stuck lane and backs the request up instead of skipping+flooding. The proactive watermark stays as the eviction target (deficit_by_owner unchanged); it is no longer a hard admission gate. The decision is computed from rank-replicated state on the collective-free reserve path, so every CP rank makes the same admit/skip choice (opus-reviewed rank-safe). Also rate-limit the four host-reservation fallback warnings (once per 10s per fallback_name, with a suppressed count) so any genuine exhaustion degrades quietly instead of producing a log storm. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
96158fa110 |
Bound overlap prefill to one pending CP batch
CP shared-KV batch planning needs the immediately previous prepared batch to remain visible as a virtual prefix, but launching another batch before processing that result can leave multiple prepared plans racing against radix insertion. The event loop now processes the previous result after planning the current batch and before launching it, preserving one-batch lookback without accumulating deeper overlap state. Constraint: CP HiCache prepared batch views are inserted during process_batch_result, not at forward launch. Rejected: Process previous result before planning current batch | loses visibility of the previous pending prepared plan needed by current planning. Confidence: medium Scope-risk: moderate Directive: Do not increase non-PP overlap depth for CP shared-KV without adding pending-radix reservation semantics. Tested: Remote cjy-glm5-new pytest test/registered/unit/disaggregation/test_overlap_disagg_prefill_event_loop.py test/registered/unit/disaggregation/test_overlap_disagg_decode_event_loop.py Not-tested: Full E2E latency impact of reduced overlap depth. (cherry picked from commit 7df8723eee8203e9e334b229178e3f8bac61396a) |
||
|
|
512fe92a83 |
Fix CP HiCache catch_up_all_layers fallback on chunked-prefill final chunk
Under overlap scheduling a chunked request's final-chunk write-backup prepare read a stale `is_chunked` (>0): that per-tick counter is decremented in process_batch_result, which the overlap loop runs AFTER the run_batch prepare hook. So prepare floored `backup_end` to a page boundary (the intermediate-chunk rule) and dropped the now-complete final tail page, while the final non-chunked insert builds the radix node at full unaligned length. The exact-equality attach predicate (prepared.logical_len == len(value)) then failed by (num_tokens-1) mod page_size, dropped the per-layer overlap backup, and forced the serial all-layer catch-up (~89% of large chunked requests in prod). Decouple the floor decision from the stale counter: the scheduler marks `req.cp_backup_is_intermediate_chunk = (req is self.chunked_req)` in `_prepare_hicache_write_backups_before_forward` (the live chunked_req identity is the authoritative "will be chunked further" signal, set this tick before run_batch), and the prepare candidate builder floors on that flag instead of `is_chunked`. Intermediate chunks still floor; only the misclassified final chunk now backs up its full tail and attaches the overlap backup. Tests: update the chunked prepare test to the new flag; add a regression test that a final chunk with stale is_chunked reserves the full tail; add an offline repro that drives the real prepare/insert/probe and sweeps unaligned tails. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2212963a6c |
Harden OpenAI tool-call/chat-template + reasoning parsing
serving_chat: - Tolerate a malformed historical tool_call `arguments` string (a valid JSON document followed by trailing content) instead of 400-ing the whole multi-turn request: salvage the leading JSON document via raw_decode, else keep the raw string. (A 112-message tool-history request was rejected with orjson "unexpected content after document".) - Catch TypeError (not only jinja2.TemplateError) from the chat-template render so a `tojson` filter on a Jinja Undefined becomes a clean 400 instead of a 500 (upstream #20700 / 5e9bd21979). reasoning_parser: - Strip only LEADING think-start marker tokens; a global replace would delete a `<think>` token that legitimately appears inside reasoning content. Preserve model-generated whitespace in reasoning/normal text (drop .strip()/.rstrip()) (upstream #24251 / dac78768f0). - Add Glm45 detector tests: leading-only strip, token-inside-content preserved, repeated leading markers, streaming trailing whitespace. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e23168e7f5 |
Fix CP shared-KV bs=1 cache-hit zero-prefix corruption; spans are sole compose input
A bs=1 cache-hit prefill produced deterministically wrong output (e.g. "0.5,0.5,0.5").
Root cause: regression
|
||
|
|
1d168d061f |
Preserve pending CP HiCache backups through final insert
Deferred non-chunked inserts can happen while a CP HiCache per-layer backup is still in flight. Rolling back the unattached prepared backup at unfinished-request time drops the only reservation that final insertion can attach, so cache_finished_req has to start a post-forward backup and catch up all layers synchronously.\n\nKeep the prepared backup for non-chunked deferred inserts and continue rolling it back for chunked middle inserts, where the suffix may be extended by later chunks and the old backup would cover the wrong range. Document the failure mode so future changes do not rediscover the same fallback path.\n\nConstraint: CP HiCache radix split cannot mutate in-flight backup nodes.\nConstraint: Chunked prefill middle inserts still need rollback because their backup range is not final.\nRejected: Always rollback unattached backups | causes post-forward catch_up_all_layers for non-chunked deferred inserts.\nRejected: Always preserve unattached backups | can attach stale backup ranges for chunked middle inserts.\nConfidence: high\nScope-risk: narrow\nDirective: Do not clear req.cp_hicache_prepared_backup on non-chunked deferred insert without proving final insert no longer needs it.\nTested: remote cjy-glm5-new py_compile radix_cache.py\nTested: remote cjy-glm5-new pytest test_cp_hicache_metadata.py::{nonchunked preserve,chunked rollback} => 2 passed\nNot-tested: live ETE replay after restarting prefill with this exact commit
|
||
|
|
ceb5345410 |
Account decode handoff queues in DP load snapshots
Decode DP dispatch was collapsing onto a few ranks because the controller only randomizes among exact minimum load pairs. The load snapshot undercounted decode handoff work: pending prefill-info requests were absent, and DecodeRequest wrappers in prealloc/transfer queues were skipped because their rid lives on .req. This makes scheduler load accounting unwrap DecodeRequest items and include pending decode requests, so TOTAL_TOKENS sees queued handoff backlog instead of repeatedly treating busy ranks as empty. Constraint: Do not mask imbalance with synthetic per-request token penalties; dispatch should be driven by accurate observed load. Rejected: Add req*4000 or other queue penalties | heuristic, workload-dependent, and hides the accounting bug. Confidence: medium Scope-risk: moderate Directive: Any new decode handoff queue must be included in get_load() or DP routing can regress to stale/min-load collapse. Tested: Remote cjy-glm5-new: PYTHONPATH=python python -m pytest -q test/registered/unit/observability/test_scheduler_metrics_load.py test/registered/unit/managers/test_prefill_adder.py -> 27 passed. Not-tested: Fresh decode ETE distribution after service restart. |
||
|
|
69ca7045ea |
Preserve chunked request affinity state
The affinity scheduler needs to know whether the current batch is led by a chunked request. The rebase carried call sites that referenced an old private field name, while PrefillAdder only retained a boolean flag, causing startup failure before scheduling could run. Constraint: CP bs>1 affinity must classify a chunked-led batch without reopening chunked-tail mixing behavior. Rejected: Recreate the old private _chunked_req_in_batch attribute | keeps a stale name and hides the public state transition in PrefillAdder. Confidence: high Scope-risk: narrow Directive: Keep chunked_req_in_batch and has_chunked_req_in_batch updated together when adding new chunked admission paths. Tested: Remote cjy-glm5-new: PYTHONPATH=python python -m pytest -q test/registered/unit/managers/test_prefill_adder.py -> 26 passed; combined run with scheduler load accounting tests -> 27 passed. Not-tested: Full ETE restart after this commit alone. |
||
|
|
ee843a946b |
Keep chunked CP prefills solo during bs>1 admission
Revert the tail-chunk co-batching gate from |
||
|
|
254d667853 |
Keep CP batch slot descriptors on IPC-compatible runtime
The syh rebase kept callers that expect reusable batch slot spans, while the restored CUDA IPC runtime lacked the helper and indexer still passed a symm-only writer-rank argument. Restore the batch-scoped span cache and remove the stale symm argument so the production compose path stays on CUDA IPC with exact per-request spans.\n\nConstraint: Production main-stream compose should use CUDA IPC, not symm writer-rank routing.\nRejected: Re-enable symm writer-rank parameters | benchmark showed no main-stream win and callers fail against IPC runtime contracts.\nConfidence: high\nScope-risk: narrow\nDirective: Keep slot-span metadata batch-scoped; do not rebuild Python span descriptors per layer without benchmarking.\nTested: Local py_compile for cp_shared_kv_runtime.py, cp_shared_kv_prefetch.py, nsa_indexer.py, nsa_backend.py.\nTested: Remote cjy-glm5-new pytest test/registered/unit/mem_cache/test_cp_shared_kv_runtime.py -> 157 passed, 2 subtests passed.\nNot-tested: Full ETE prefill/decode runtime after restarting services. |
||
|
|
4d5c7f32d6 |
Keep CP shared-KV prefetch warnings actionable
Expected no-prefetch paths were polluting production logs: no cache prefix, tiny/first-layer windows, and FP8 RAGGED top-k were being reported as fallback warnings. The prefetch contract now treats zero-prefix and first-layer misses as normal skips, while preserving warnings for non-zero misaligned prefixes and real consume misses after the first layer. The same change keeps RAGGED cache-hit prefetch eligible and records the CE/IPM prefetch contract in the plan doc. Constraint: FP8 sparse prefill uses RAGGED top-k, but CP shared-KV prefix materialization is still page-slot based Constraint: Layer 0 has no previous attention-window hook that can have prefetched the layer Rejected: Warn whenever a prefetcher is absent | no-cache and too-short requests are expected synchronous paths and make logs unusable Confidence: high Scope-risk: moderate Directive: Keep CP_SHARED_KV_FALLBACK warnings for unexpected contract failures only; use debug logs for expected skip paths Tested: Local py_compile for cp_shared_kv_prefetch.py, nsa_indexer.py, nsa_backend.py Tested: Remote cjy-glm5-new targeted regression: 3 passed, 21 warnings Tested: Remote cjy-glm5-new full test_cp_shared_kv_runtime.py: 156 passed, 21 warnings, 2 subtests passed Not-tested: New ETE run after prefill restart to confirm log volume reduction in production traffic (cherry picked from commit e08e321e5929fdbb30102ec0b19c6ff0ecac7e7e) |
||
|
|
cd4412a4b8 |
Reuse IPC descriptors across CP shared-KV layers
CP shared-KV slot remaps already have forward-batch lifetime, but the IPC materialize path rebuilt owner/source/dense descriptor tensors on every layer. Cache prefix/current IPC descriptors on the token and paged slot-remap objects, keyed by layout, spans, device, descriptor kind, and prefix capacity, so all model layers can reuse the same request/batch-plan descriptors. Constraint: Small-extend cache-hit workloads showed descriptor setup could exceed the all-reduce baseline before any IPC kernel work ran. Rejected: Global descriptor cache | slot-remap lifetime is safer and avoids stale entries across request/batch-plan changes. Rejected: Cache without physical page capacity | prefix descriptors encode capacity-invalid pages and must miss when capacity changes. Confidence: high Scope-risk: moderate Directive: Do not reuse descriptors across different slot_logical_pages identity, CP layout, spans, device, or prefix capacity; stale descriptors can alias dense slots across requests. Tested: Local py_compile; local git diff --check; remote g0034 cjy-glm5-new targeted descriptor tests 2 passed; remote full test_cp_shared_kv_runtime.py 146 passed, 21 warnings, 2 subtests passed. Not-tested: Full ETE throughput/accuracy after descriptor cache; CUDA service benchmark still needed to quantify speedup. (cherry picked from commit addd1ca1571e41458315d15304a0e841682fe8fa) |
||
|
|
bafb55044b |
Keep CP current IPC staging proportional to touched pages
Cache-hit bs>1 current reuse can create very large dense attention buffers while touching only a small set of current pages. The previous SGLang runtime asked tai-kernel for a staging buffer sized like the full dense tensor, which caused CUDA OOM before the current IPC fast path could run. Switch token and index current IPC helpers to descriptor-compact staging: publish the dense destination pages into compact staging slots and materialize peers from compact source page ids back to the original dense destination pages. Document the failure mode and the compact-staging contract so the dense-sized contract is not reintroduced. Constraint: CUDA + SGLANG_CP_SHARED_KV_USE_TAI_MATERIALIZE=1 must fail fast instead of silently falling back to current-slot all_reduce Rejected: Let staging allocation failure fall back to all_reduce | hides the bug and restores the expensive collective path Rejected: Size staging by the full dense tensor | reproduces the 965MB staging OOM on long-prefix cache-hit batches Confidence: high Scope-risk: moderate Directive: Current IPC helper source ids are compact staging ids; destination ids remain dense slot pages Tested: Remote cjy-glm5-new PYTHONPATH=python:/mnt/beegfs/cjy/tai-kernel/python python -m pytest -q test/registered/unit/mem_cache/test_cp_shared_kv_runtime.py -> 144 passed, 2 subtests passed Tested: Local py_compile cp_shared_kv_runtime.py Not-tested: Full ETE service restart with production traffic after this commit (cherry picked from commit 906ecbe5d4f08b73242e98e2b628e26516d5b04a) |
||
|
|
142e7a5a64 |
Keep CP shared-KV fast paths off dense current collectives
CP shared-KV cache-hit batches should compose long prefix pages and short current pages through page-slot IPC instead of falling back to dense all_reduce. Wire the runtime and prefetch consume paths to the TAI current-staging helpers, fail fast when the configured CUDA fast path cannot run, and document the bs>1 cache-hit benchmark evidence. Constraint: bs>1 prefill must preserve the page-slot contract across fp8/bf16 and zero-lane current tails. Rejected: Silent all_reduce fallback | hides correctness and performance regressions in production. Confidence: medium Scope-risk: moderate Directive: Any future fallback in CP shared-KV CUDA fast paths must be explicit warning/fail-fast and covered by runtime tests. Tested: Local py_compile cp_shared_kv_runtime.py and cp_shared_kv_prefetch.py; remote PYTHONPATH=python pytest -q test/registered/unit/mem_cache/test_cp_shared_kv_runtime.py (144 passed, 21 warnings, 2 subtests passed); remote TAI IPC benchmark fp8 bs>1 cache-hit matrix recorded in docs. Not-tested: Full ETE mixed replay after replacing all current collectives with IPC. (cherry picked from commit 8aa3b4ce59e0ebef5da5b0d07499a5f1d9785997) |
||
|
|
2387787ebc |
Fail fast when CP compose would hide dense fallback
CP shared-KV bs>1 compose must not silently fall back to dense full-buffer collectives when CUDA TAI materialize is expected. The fallback masks both correctness contract drift and severe synchronization/communication regressions, especially while comparing the symm path with the older IPC path.\n\nThis keeps CPU/unit-test fallback available, but makes production CUDA+TAI runs raise an explicit compose_v2 fail-fast for token-KV and index dense fallback. It also records the symm-vs-IPC comparison contract so barrier and collective counts are evaluated alongside elapsed time.\n\nConstraint: Production cache-hit-heavy bs>1 paths must expose unexpected dense collectives instead of silently taking them.\nRejected: Cherry-pick the old IPC branch wholesale | it conflicts with the symm compose design and would mix two transport protocols before benchmarking.\nRejected: Allow dense fallback with warning only | warning can be missed and still corrupts performance conclusions.\nConfidence: medium\nScope-risk: moderate\nDirective: Do not re-enable dense full fallback in CUDA+TAI compose paths without a benchmark proving it is intentional and a correctness test covering cache-hit bs>1.\nTested: python -m py_compile for cp_shared_kv_runtime.py and test_cp_shared_kv_runtime.py; git diff --check.\nNot-tested: Remote container pytest/ETE; local pytest is not reliable in this workspace because dependencies such as orjson are missing. |
||
|
|
9d65bdba95 |
Measure mqa-logits chunk pipelining: not worth it (S2b closed)
The bench gains a pipelined mode (topk_N event-gated on a side stream,
overlapping logits_{N+1}). Byte-equality validation was dropped after
establishing the kernel chain is run-to-run nondeterministic even
serial-vs-serial (near-equal fp32 selection).
Result on g0033 H200: pipelining recovers only 0.3-8.9% where small
chunks cost +15-46% — fp8_mqa_logits and fast_topk_transform_fused are
both SM-saturating, so concurrent streams timeshare instead of
overlapping; the small-chunk penalty is small-M GEMM inefficiency. No
production pipeline path; the serial loop at CHUNK_MAX_GB=2 stands.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
6a1e862f48 |
Group cache-hit prefills into dense batches (SGLANG_CP_PREFILL_AFFINITY_GROUP)
专题 S4 (design: docs_internal/perf/prefill-compute-intensity-plan.md S4, amended). Under FCFS a cold request joining a warm-led batch turns a 1-2s cache-hit forward into a 5-10s one, splitting the warm work into the 新-cache-新 pattern. The policy prevents exactly that one thing: - WARM candidates always admit (into a cold-led batch they are free density — the cold extend dominates the forward anyway). - COLD admits into an empty or cold-led batch (small colds co-batch today; the FCFS head always starts a batch so the queue keeps moving). - COLD into a WARM-led batch is skipped, bounded by a per-pass window (W=16 skips), a head defer count (K=3 passes) and an age bound (T=5s). On any bound the scan STOPS instead of force-admitting: the cold waits for the same forward either way, but leads its own clean batch next pass instead of polluting this one. The skip is strictly post-match / pre-admit (after init_next_round_input, before add_one_req): no lock, no allocation, no budget mutation to unwind, and re-matching a skipped candidate next pass is exactly what the scan already does after a cap rejection. Classification is the in-scan match result (device prefix + host hit vs a 64-token floor) — under FCFS+L2 no pre-scan signal exists, so this adds zero matching work for inspected candidates. Disabled wholesale under priority scheduling (the skip must not reorder across priority classes). Three amendments vs the design draft, reasoned in the decision-table docstring: cold+cold-led admits (STOP would regress today's small-cold co-batching); starved heads STOP rather than force-admit (clean batch boundaries at identical latency); priority interaction handled by disabling rather than per-request comparison. Decision logic is a pure function with table + bounds unit tests (28/28 adder suite green). Default OFF. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
54c056af83 |
Let tail chunks co-batch with cache-hit requests (SGLANG_CP_PREFILL_MIX_CHUNKED)
专题 S1 (design: docs_internal/perf/prefill-compute-intensity-plan.md S1.0-S1.11). A batch containing a chunked prefill request has been forced to bs=1 by the CP gate, so every chunk of a long prompt monopolizes a forward while short-extend cache-hit continuations queue — the direct cause of the replay TTFT tail (p90 19.3s / p99 49.2s at 91.8% cache hit). Yet mixed chunk batches already occur today (a freshly-chunked request keeps earlier-admitted small requests), proving the CP forward path is mixed-chunk-safe; only admission was asymmetric. Three changes, the first flag-independent: - add_chunked_req now seeds the budget with the chunk's TRUE prefix (was 0), so the CP cached tally and the buffer estimator's mqa_logits k_rows see the chunk's footprint before any later request is admitted (landmine D1). - New SGLANG_CP_PREFILL_MIX_CHUNKED (default OFF): with it on, the gate admits requests after a chunked one and lets the existing CP caps (extend / cached / buffer, now correctly seeded) bound the batch — a FULL chunk still ends the scan by consuming the chunk-clamped extend cap; only a tail chunk leaves headroom. A chunked prefix that is not page-aligned (rare sub-page final-chunk tail) keeps its batch solo (the CP page-aligned split would fail-fast otherwise). - The symm staging capacity identity (admission extend cap + request slack == staging pages) is asserted when the flag is on, locking the coupling the design relies on (plan doc S1.4 I2). Tests: 4 new adder units (budget seeding; tail chunk admits followers; full chunk solo by budget; non-aligned prefix solo); the 8-rank byte-exactness scenario gains a chunk-shaped request (2048-token page-aligned carried prefix + 512 extend) — all four phases (legacy/v2/symm/prefetch) byte-identical on g0033. Known pre-existing cross-file pollution noted in problems.md P16. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f9ce28ee5c |
Add mqa-logits chunking cost micro-benchmark (专题 S2a)
Measures the real cost of shrinking SGLANG_NSA_MQA_LOGITS_CHUNK_MAX_GB: the faithful indexer loop (deep_gemm.fp8_mqa_logits + fast_topk_transform_fused, serial) at GLM-5.1 shapes (H=32, D=128, topk=2048) across cold-chunk / tail-chunk / warm-continuation / warm-long scenarios. g0033 1xH200 results: 2GB costs at most +5.5% (cold 64K chunk) and is -6.7% on the heaviest warm-long shape; the knee is ~1GB; 0.5GB is +30%. Shrinking 8->2GB frees ~6GB of the per-batch CP admission budget for KV layer buffers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
58f2350738 |
Give prefill inflight transfers a liveness bound
E2e caught one request wedged FOREVER in disagg_prefill_inflight_queue (probes 21s apart with zero traffic both showed #inflight-req: 1, no reap/timeout warnings ever logged). Mechanism, established by reading the full state machine: prefill Success is set locally by the transfer worker on the LAST chunk; if the decode peer is torn down between the handshake and the prefill's final send(), add_transfer_request silently drops the chunk (no transfer destinations) — Success becomes unreachable. The only external rescue, the decode ABORT notification, is best-effort (silently swallowed on send error, no-op if it races the room registration), there is no prefill-side heartbeat of decode sessions, and the sender's only timeout covers Bootstrapping — the inflight queue itself has no liveness bound. The orphan pins the request's KV pages and rides every poll collective. Two fixes, both reaped through the existing Failed branch via the CP/TP MIN-reduce poll consensus (Failed=0 wins, so one rank concluding flips every rank together — rank-uniform by construction): - add_transfer_request: a room with no transfer destinations that is NOT already Success (the dummy-rank handshake marking) now concludes Failed loudly instead of dropping the chunk silently. - Inflight residency timeout: entries stuck in a non-terminal poll state past SGLANG_DISAGGREGATION_INFLIGHT_TIMEOUT (default 300s, matching the sibling BOOTSTRAP/WAITING timeouts) get sender.abort() and reap on the next poll. Covers what the hardening cannot: lost ABORT datagrams, decode crashes. Known sibling gaps left for follow-up: the decode transfer queue has no Transferring liveness bound, and an abort that matches no queue is still a silent no-op (much narrower race than first thought — work requests are ordered before control requests within a tick). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d66758d929 |
Run the symm staging fill on byte views (no fp8 index_copy kernel)
First e2e launch crashed every CP rank at the symm token fill: index_copy_cuda is not implemented for Float8_e4m3fn — the production KV pool dtype, which the 8-rank test missed by building its pools as uint8. The fill is a whole-token-row copy, so it is dtype-agnostic: both the staging span and the current rows now go through uint8 views. The 8-rank test now builds the KV pools and current rows as float8_e4m3fn (payloads constructed as bytes, compared as bytes) so the production dtype is what every phase exercises. Validated on g0033: all four 8-rank byte-exactness phases under fp8. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1923fcd67e |
Wire the symm current exchange into the CP shared-KV prefetchers
The bs=1 MLA/index prefetchers replaced their consume-side trailing-range NCCL all-reduce with the staging exchange: fill current rows straight into this round's staging span (token-KV collapses to one cached index_copy; the index fill kernel is just pointed at the staging page inverse), cp_symm_barrier, then gather ALL current pages — this rank's own included — from the stagings into the prefetched dense buffer. The symm+prefetcher FAIL_FAST is gone. Rank-uniformity moves with it: staging registration now also happens in maybe_create (batch-logical gates, before any per-rank miss can diverge), because with a prefetcher active the sync compose runs only on per-rank misses and its lazy collective registration would hang. A hit/miss divergence itself stays barrier-safe — both the prefetch consume and the sync-compose fallback execute exactly one begin_round + barrier per (layer, kind), and the counting barrier is shape-free (unlike the AR pair it replaces, which would shape-mismatch). Found by the new index test phase: the fill/remap kernel family skips page id 0 as the SGLang dummy page, so a 0-based first staging slot was never written. The staging layout now reserves row 0 (slot of current page i = i + 1) for every kind, matching the convention instead of depending on per-kernel behavior. Launch-path cost: per-(kind,parity) peer pointer tables and the [pool|staging] concatenations are precomputed/cached (identity pinned by holding the pool-table reference); all prefetch descriptors, staging row indices, and mixed_locs are built once per batch. Validated on g0033 8xH200: 151 unit tests; 8-rank byte-exactness for token sync symm (8 layers), index sync symm (4 layers, new phase), and MLA + index prefetch consume_prefix_with_current vs the legacy sync compose. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d63fbd4d79 |
Fold the symm compose into one barrier-gated mega gather
The publish-variant staging exchange measured ~equal to the default compact-current AR (90.4 vs 86.5 ms/batch on the traced scenario): the publish copy and the barrier serialized behind the 0.65 ms prefix gather ate the transport win the isolated current exchange showed (0.196 vs 0.354 ms). Fix the structure instead of the copy: current rows are now written straight INTO the staging — the fill kernels take their write destinations solely from page_inverse, so a per-batch staging-remapped page inverse on the plan retargets them with zero kernel changes — then cp_symm_barrier, then ONE slot-dense gather covers prefix pages (pool pointers) and ALL current pages (staging pointers, including this rank's own) through a concatenated 2*cp pointer table where current slots carry owner = cp_size + writer and src = staging slot. No publish copy, no prefix pre-gather, no second gather. The fused fill's loc outputs are dense-geometry-bound, so the token-KV path computes mixed_locs/staging row indices once per batch (they are layer-invariant) and the per-layer fill collapses to a single index_copy_ into the zeroed staging span. Benchmark (g0033 8xH200, byte-exact, idle-checked): 62.8 ms/batch vs 84.4 default Step A (-26%) and 60.8 ideal; publish variant was 88.0. 151 unit tests; 8-rank GPU byte-exactness vs v2 across 8 layers, arena on and off. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e8e5edb230 |
Shrink the symm compose region to a compact current-page staging
Peers only ever read the CURRENT pages of a rank's compose output — the prefix comes straight from the IPC-registered KV pool — so the symm region does not need to hold the whole dense buffer (pool-bound ~2.5 GB double-buffered slab). It now holds one round of current pages in merged-span order (extend-cap-bound, ~58-100 MB), and dense buffers become purely rank-local (plain allocations or the optional local arena; COMPOSE_SYMM no longer requires COMPOSE_ARENA). Exchange per compose call: publish my written current pages dense[page] -> staging[slot i] (slot = the page's batch current index, identical on every rank, so peers address each other's staging with no per-batch handshake), cp_symm_barrier, gather peers' staging[writer][slot] -> dense[page] via the existing src!=dst page gather. Reuse safety keeps the parity-half distance-2 argument, now on the staging. Capacity sizing comes from the admission caps (max_total_extend_tokens / max_batch_requests) with a pool-derived fallback and the SYMM_HEAP_MB override; overflow fails fast (batch-logical, hence rank-uniform). Idea credit: laoyao0822's touched-pages-proportional staging (906ecbe5d4), rebound onto our barrier-gated, group-agreed transport. Validated on g0033 8xH200: 151 unit tests; 8-rank GPU byte-exactness vs compose_v2 across 8 layers (arena on and off, parity halves exercised); benchmark path e (real protocol) byte-exact, current-page exchange 0.196 ms vs 0.354 ms compact-AR isolated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f74f0a557f |
Keep CP pending-split probes collective-free
Radix insert can reach the pending-backup/stale-tail split probe on only a subset of CP ranks. The previous opportunistic ack drain called writing_check(), which may issue a scalar write-visibility collective while other ranks are polling inflight transfer state with a different tensor shape. That ordering mismatch matches the observed Gloo 4-vs-2 abort after CP_HICACHE_FALLBACK insert_deferred_pending_backup_split logs. This keeps the split probe local and conservative: it leaves ready write acks queued and lets the globally ordered scheduler/cache-event path drain them. The regression test locks the contract by making any collective from the pending-split probe fail. Constraint: CP radix insert/pending-split probing is not globally ordered across ranks. Rejected: Drain ready write acks from the split probe | can interleave with scheduler poll collectives and corrupt distributed collective ordering. Confidence: high Scope-risk: narrow Directive: Do not call writing_check() or any CP/TP collective from radix pending-split probes; drain write acks only from globally ordered paths. Tested: Remote cjy container: PYTHONPATH=python python -m pytest -q test/registered/unit/mem_cache/test_cp_hicache_metadata.py => 126 passed, 21 warnings. Tested: Remote cjy container: python -m py_compile python/sglang/srt/mem_cache/hiradix_cache.py => PY_COMPILE_OK. Not-tested: Full ETE prefill/decode replay after this change. Co-authored-by: OmX <omx@oh-my-codex.dev> |
||
|
|
a24111a5f4 |
Cut per-layer CPU on the prefill launch path: validators, plans, spans
From the nsys CPU-gap attribution (launch thread, one 78-layer forward: 374ms API time; 642 cudaStreamSynchronize blocking 89.5ms and overlapping 122ms of the 505ms GPU idle; ~44ms pure-Python before concat_mla_absorb_q): - memory_pool_host: skip validate_page_aligned_token_indices on CUDA tensors in _get_indexer_page_indices and _prepare_load_page_indices — torch.any/torch.equal there cost a queue-deep cudaStreamSynchronize per layer-group submit (~0.42ms each, ~12.7ms/forward measured). Same construction-based-invariant guard the CacheController pair check already documents; CPU/test tensors stay validated. - nsa_indexer: per-batch _CpRaggedIndexPlan replaces the per-F-layer rebuild of the O(total-q-tokens) topk offset list and the 6-7 int32 ragged descriptor tensors (segment records, kv_lens/q_starts/q_lens/ k_bases/q_bases/current_bases). All inputs are batch metadata; the plan is anchored on the forward batch with a content key over cp_index. - nsa_indexer forward_indexer: read seq_lens_cpu instead of a device seq_lens[i].item() per request per layer (one stream sync each). - cp_shared_kv_runtime: get_or_build_batch_slot_spans caches the layer-invariant prefix/current slot spans per batch (the builders read logical_pages only for its shape); nsa_backend x3 + nsa_indexer call sites switched. Microbenchmark (idle H200, traced batch shape bs=12 / 44.6K q tokens, test/manual/bench_cpu_gap_fixes.py, equality-checked): validator path 197.1us -> 59.2us per submit under a busy queue (x3.3); ragged plan 3238.6us -> 36.2us per layer (x90, ~128ms launch-thread time per forward at 40 F-layers); slot spans 20.1us -> 0.5us (x41). Layer suites A/B vs HEAD: identical failure set (5 pre-existing CPU-tensor indexer tests), no regressions. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b356773d2f |
Harden symm compose per review: prefetch gate, round parity, hot-path CPU
Addresses the opus review of the Step A/B compose stack: - FAIL FAST on symm + prefetcher coexistence: the MLA/index prefetchers issue their own per-span collectives and bypass the symm exchange, so letting them coexist would make COMPOSE_SYMM a silent no-op on every prefetch hit (measured "parity" that never ran). maybe_build_current_ page_writer_ranks now raises when a prefetcher is attached or SGLANG_CP_SHARED_KV_ENABLE_MLA_PREFETCH is set. - Arena parity now derives from a monotonic round epoch instead of raw layer_id: EAGLE draft layers reuse decoder layer ids, so id-parity would stop alternating halves (draft L0 -> next forward target L0 lands on the same half) and repeated-id rounds would keep appending into one half. begin_round(layer_id, kind) starts a new round (other half, offsets reset) when the id changes OR a kind repeats; regression test included. - Per-batch writer-ranks cache on the forward batch and a presence flag in the plan key: previously the ~bs x current-pages writer list was rebuilt AND tuple-hashed on every layer per call site, pure launch-path waste. - Rank-uniformity invariant documented at _symm_exchange_current_pages (any per-rank gate must go through a group agreement first) and an actionable arena-overflow message naming SGLANG_CP_SHARED_KV_SYMM_HEAP_MB. Deferred follow-ups (documented in the design doc): factor the two near- clone v2 helpers, single barrier per F-layer (~308us/batch), persistent device-side peer-ptr tables for the gather. 182 targeted tests + 8-rank GPU byte-exactness (v2 and symm) re-validated on g0034 syh-dev-new. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c55e406176 |
Add symm-heap current-page exchange for CP shared-KV compose (Step B)
Replaces the last remaining NCCL collective in the bs>1 compose layer loop (the compact-current all-reduce) with a DeepEP-style counting barrier plus a CUDA-IPC gather from peers' symmetric dense buffers, behind SGLANG_CP_SHARED_KV_COMPOSE_SYMM (+ COMPOSE_ARENA, both default off). - CpComposeArena.register_symm: fixes capacity at the pool-derived bound (logical pages x dense page unit, overridable via SGLANG_CP_SHARED_KV_SYMM_HEAP_MB), allocates slab + flags in CUDA-IPC memory, exchanges handles once over the CP group; deterministic bump carve means peer_base + my_offset addresses any peer's dense buffer with no per-layer handshakes. Registration happens only in the token-KV compose (uniform first-use point); growth after registration raises. - Current pages are single-writer at page granularity under the page-aligned in-seq split, so the exchange is the existing gather_cuda_ipc_peer_pages with src==dst page ids and writer (compute owner) ranks; writers are built per batch by build_batch_current_page_writer_ranks and gated by maybe_build_current_page_writer_ranks (page_aligned metadata required). The barrier runs even with zero remote pages (counts must match). - _agreed_tai_ipc_peer_ptrs: the per-rank IPC capability probe is now agreed across the CP group (one-time MIN all-reduce per pool tensor) so ranks can never split between gather and collective paths and deadlock on mismatched NCCL shapes. - ComposePlan cache re-anchored ON the slot_remap object (forward-batch lifetime) instead of a module-level tensor-identity key, which could go stale when a freed tensor's address is reused by the next batch. Validation (g0034 syh-dev-new): tai-kernel cp_symm_barrier correctness (200 adversarial iterations, rotating 10ms producer delays, phase-safety, flags drained at quiescence) and perf (7.7us max-rank latency) both pass; 8-rank GPU test extends to the symm path - byte-identical to compose_v2 across 4 layers (parity halves exercised, slab registered); mem_cache suite 464 passed with the only failures being a documented pre-existing sys.modules stub pollution pair, reproduced identically with SGLANG_CP_SHARED_KV_COMPOSE_V2=0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f47b739e30 |
Restructure bs>1 CP shared-KV compose to one gather + one collective
The bs>1 partial-current materialize issued one sum-all-reduce per request span per buffer per layer (48 collectives per F-layer at bs=12, 2880 per batch, 419ms + 49ms launch gaps in the production trace), all inline on the compute stream. The data is a partitioned gather, not a reduction: every byte has exactly one producer. compose_v2 (SGLANG_CP_SHARED_KV_COMPOSE_V2, default on) replaces this with: - fast path: one tai-kernel CUDA-IPC slot-dense gather covering ALL prefix spans (full-range descriptors, -1 sentinels zero-fill current slots and replace the dense zero-fill) + ONE collective over the compact current pages (uint8 byte view; exact because every byte is writer-exclusive). - fallback (no peer IPC): local materialize of all prefix spans + ONE whole-buffer sum-all-reduce (rows are still writer-exclusive pre-reduce). IPC capability is decided once by the cached peer-pointer probe; after a successful probe a failing gather raises (no per-call try/except). cp_shared_kv_compose.py adds the per-batch ComposePlan descriptor cache (layer-invariant, keyed on the slot_logical_pages identity) and the CpComposeArena with tier-S carve discipline (deterministic bump, layer- parity halves; default off) so the Step B symmetric-memory conversion is a registration flip. Microbenchmark (g0034 8xH200, traced 12-req batch, per batch): per-span 214ms -> fused AR 119ms -> IPC prefix + compact current 84ms; symm target 61ms. Validation: 143 unit tests incl. v2-contract twins, legacy siblings and rank-merged simulations under both paths; mem_cache dir 432 passed; 8-rank GPU byte-exactness vs legacy with real NCCL + IPC (test/manual/ test_cp_shared_kv_compose_v2_8rank.py) passed with no fallback markers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ecf7998c80 |
Route undrained-ack reservation refusal around the capacity retry
The reserve_write_cp undrained-ack refusal returned a bare HiCacheWriteFailure, which _reserve_write_cp_indices_no_collective interpreted as a host-capacity failure: with free host space the retry path tripped the predicted-no-deficit RuntimeError (crashing the scheduler in exactly the scenario the gate exists to handle gracefully), and with a deficit it triggered pointless evictions before refusing again. Give HiCacheWriteFailure an explicit reason (default host_capacity keeps all existing constructors/semantics); the wrapper returns non-capacity failures immediately as skip-this-round, and a wrapper-level test pins that undrained_ack reaches neither the retry admission nor host eviction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
47cb9b9d11 |
Make the duplicate-reservation ack gate O(1) for fresh node ids
node_has_undrained_write_ack scanned ack_write_queue on every reservation, including the per-request-per-chunk prepare path whose node ids are freshly minted and can never be queued. Track the max node id ever appended to the queue: any queued id is <= max by construction (no monotonicity assumption needed for correctness), so fresh ids exit on one integer compare and the scan remains only for re-reservations of old ids (the rare write_backup fallback). All three ack_write_queue append sites now go through a single _append_write_ack funnel that maintains the max — the zero-owned-rank and non-CP write acks previously would have bypassed it, allowing false negatives on ranks owning no pages of a node. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
698a3e2431 |
Enforce one in-flight CP write ack per node at the producer
|
||
|
|
dfa168abe9 |
Prevent duplicate CP HiCache write acks from crashing split drain
CP per-layer write-through can leave duplicated ready ack ids in the ack queue while pending-split insertion drains write visibility. The pre-scan observed both duplicates as registered before the first completion removed radix state, so the second completion raised KeyError and killed the scheduler. The write-check path now records completed ack ids until matching queue entries are drained. Duplicate completed acks are removed with a warning, while genuinely unknown or unregistered acks still fail fast. Constraint: CP HiCache pending-split drain may call writing_check outside the normal idle event path. Rejected: Blind pop(..., None) for all unknown acks | would silently hide truly unattached or corrupt ack state. Confidence: high Scope-risk: moderate Directive: Do not remove the completed-ack short-term memory unless duplicate and stale ack queues are proven impossible across TP MIN synchronization. Tested: Local py_compile for hiradix_cache.py and test_cp_hicache_metadata.py. Tested: Remote cjy-glm5-new py_compile and full test_cp_hicache_metadata.py, 121 passed. Not-tested: Full ETE prefill workload after restarting service. Co-authored-by: OmX <omx@oh-my-codex.dev> |
||
|
|
75d7d8772e |
Unify over-length errors into the PayloadTooLargeError 413 format
Over-long inputs produced two different client errors depending on
which bound rejected them: the TokenizerManager pre-check (raw
context_len) returned 413 PayloadTooLargeError ('The input (N tokens)
is longer than the model's context length (M tokens).'), while inputs
between that and the scheduler's stricter effective limit hit
validate_input_length and returned 400 BAD_REQUEST with different
wording (and a confusing 'X exceeds X' message since the check is >=).
Unify on the 413 format end to end:
- validate_input_length wording now matches the TokenizerManager
message, reporting the effective per-request limit.
- set_finish_with_abort takes status_code/err_type; the scheduler
length-rejection sites abort with REQUEST_ENTITY_TOO_LARGE +
PayloadTooLargeError. The batch handler previously queued the
over-long request WITHOUT marking it aborted (it proceeded to
prefill) — also fixed.
- Non-streaming aborts with 413 raise PayloadTooLargeError (now a
ValueError subclass so raw /generate-style endpoints that only
catch ValueError still respond; the OpenAI layer's except clause
is reordered to win and emit the 413 format).
- Streaming abort responses prefer the scheduler-provided err_type
over the HTTPStatus name.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
9c8e3e99cb |
Align OpenAI serving behavior with Para deployments
Absorb PR 11's final Para compatibility surface as an opt-in OpenAI serving layer rather than hard-coding business defaults into protocol models. The change adds server args for Para chat defaults, Kimi/GLM compatibility, tool-choice normalization, tool-role text flattening, and streaming first-chunk error preflight while preserving default upstream behavior unless explicitly enabled. Reasoning token usage is also propagated through chat/completion usage paths, with GLM compatibility emitting completion_tokens_details.reasoning_tokens. Low-risk protocol fixes accept string image_url content parts and preserve GLM function-call argument value whitespace. Constraint: Online Para-compatible deployments require request/response semantics that differ from default OpenAI serving behavior. Constraint: Current CP/HiCache/bs>1 work must not be coupled to OpenAI serving compatibility changes. Rejected: Merge PR 11 history directly | intermediate commits briefly hard-code chat max_tokens=32768 before later gating it by server args. Rejected: Enable Para compatibility by default | would change non-Para OpenAI-compatible deployments. Confidence: high Scope-risk: moderate Directive: Keep Para-specific serving policies behind explicit server args unless the business contract changes globally. Tested: PYTHONPATH=python:. python -m unittest discover -s test/registered/unit/entrypoints/openai -p 'test_para_serving_protocol.py' -v (19 tests OK) Tested: python -m py_compile modified OpenAI serving, tokenizer manager, server_args, function-call detector, and test files Not-tested: Live router/prefill/decode OpenAI serving E2E after enabling Para flags. Co-authored-by: OmX <omx@oh-my-codex.dev> |
||
|
|
8cfdb0466e |
Reduce CP per-layer transfer success-log noise
Per-layer prefill-to-decode transfer now finishes per request/rank/chunk, so success-path INFO logs can dominate production logs and hide actual failures. Keep successful finish breakdown and completion messages at DEBUG while preserving nonzero finish status as WARNING. Constraint: Per-layer transfer is a hot path under CP shared-KV and may produce many batch completions per request. Rejected: Disable CP per-layer transfer logging entirely | failures still need visible warning-level evidence. Confidence: high Scope-risk: narrow Directive: Do not promote successful per-request transfer completion logs back to INFO without rate limiting. Tested: PYTHONPATH=python python -m pytest -q test/registered/unit/disaggregation/test_cp_per_layer_transfer.py::TestPerLayerTransferContext::test_successful_finish_does_not_emit_hot_path_info_log Tested: python -m py_compile python/sglang/srt/disaggregation/cp_per_layer_transfer.py python/sglang/srt/disaggregation/mooncake/conn.py Not-tested: Full local disaggregation suite blocked by missing local orjson dependency. Co-authored-by: OmX <omx@oh-my-codex.dev> |
||
|
|
7284a469a2 |
Reuse prepared HiCache load descriptors across CP prefill layers
CP shared-KV bs>1 cache-hit loads already merge request load ops, but the host pool still rebuilt layer-invariant mapping work from the same host/device indices. Introduce a PreparedLoadDescriptor lifecycle around begin/end load, wire MLA KV and NSA index H2D loads through tai-kernel prepared submit when available, and add timing hooks plus regression coverage for descriptor reuse and explicit fallback logging. Record the P4/P6b design and benchmark results in the advanced feature notes. Constraint: Radix residency and allocator decisions remain synchronous; only the data-transfer descriptor is prepared for per-layer async submit. Constraint: Production fast path must not silently fall back when tai prepared H2D support is missing. Rejected: Cross-batch descriptor reuse | descriptor lifetime and tensor ownership are only safe within one load operation. Rejected: Change L2->L1 scheduling to layer-ahead prefetch in this commit | that is a separate lifecycle change after descriptor reuse is stable. Confidence: medium Scope-risk: moderate Directive: Keep LayerDoneCounter per-layer readiness semantics; do not replace with all-layer waits. Tested: python -m py_compile python/sglang/srt/mem_cache/memory_pool_host.py python/sglang/srt/managers/cache_controller.py Tested: Remote g0034:cjy-glm5-new PYTHONPATH=python python -m pytest -q test/registered/unit/managers/test_hicache_controller_cp.py (88 passed) Tested: Remote tai-kernel prepared descriptor CUDA test (6 passed) and P4 benchmark full matrix (90 rows) Not-tested: ETE replay/GSM8K cache-hit correctness after this commit Not-tested: Layer-ahead L2->L1 prefetch scheduling Co-authored-by: OmX <omx@oh-my-codex.dev> |
||
|
|
adf357b02c |
Default CP bs>1 extend admission to chunk budget
When chunked prefill is active, CP shared-KV bs>1 cannot consume more extend tokens than the current chunk budget. If the CP-specific extend-token limit is omitted, default it to rem_chunk_tokens so scheduler admission reflects the reachable chunk capacity. The request-count and cached-token knobs keep their None-as-unlimited behavior. Constraint: CP bs>1 batching must not advertise a larger extend batch than chunked prefill can execute. Rejected: Require users to always set --cp-shared-kv-prefill-max-total-extend-tokens | the safe default is already available from chunked prefill state. Rejected: Default batch request or cached-token limits | those are policy knobs and None should remain unlimited. Confidence: high Scope-risk: narrow Directive: Keep --cp-shared-kv-prefill-max-total-extend-tokens as min(user_limit, chunk_budget) when both exist. Tested: Local py_compile for schedule_policy.py and test_prefill_adder.py. Tested: Remote g0034 cjy-glm5-new targeted prefill_adder tests: 2 passed. Not-tested: Full ETE scheduler batching distribution after defaulting the extend limit. Co-authored-by: OmX <omx@oh-my-codex.dev> |
||
|
|
d696039092 |
Align MQA logits admission with CP split runtime shape
CP shared-KV batching previously estimated MQA logits from full request extend/context rows, which overstated memory because CP in-seq split only computes each rank's two zigzag segments. Add CP-size aware row accounting that mirrors the fused CP MQA materialization path and take the worst local rank peak for scheduler admission. Expose SGLANG_NSA_MQA_LOGITS_CHUNK_MAX_GB as a more direct cap for one fp32 MQA logits chunk. Runtime and scheduler now both translate this GB cap into chunk rows from the actual K rows, while keeping the old row cap as a mutually-exclusive expert override. Constraint: Scheduler admission must stay CUDA-sync-free and use static budget information only. Rejected: Keep full-request q*k admission | it over-gates CP bs>1 batches because CP splits q rows per rank. Rejected: Let rows and GB caps both apply | precedence would be ambiguous during tuning. Confidence: medium Scope-risk: moderate Directive: Keep MQA logits admission tied to the fused CP MQA segment shape; do not revert to full request token counts. Tested: Local py_compile for touched runtime, scheduler, estimator, and tests. Tested: Local pytest test_cp_shared_kv_prefill_buffer_estimator.py: 8 passed. Tested: Remote g0034 cjy-glm5-new py_compile and targeted estimator/runtime tests: 13 passed. Not-tested: Full ETE high-cache-hit CP bs>1 load with SGLANG_NSA_MQA_LOGITS_CHUNK_MAX_GB. Co-authored-by: OmX <omx@oh-my-codex.dev> |
||
|
|
250fab291d |
Account for MQA logits in CP batch admission
CP shared-KV bs>1 admission already bounds request count, extend tokens, cached tokens, and an estimated temporary buffer size. The estimate missed the fp32 MQA logits temporary, whose peak grows with query rows times context rows and can dominate high-cache-hit multi-request batches. Add an MQA logits peak term to the CPU-only estimator and include it in the layer-forward peak enforced by --cp-shared-kv-prefill-max-buffer-size. When SGLANG_NSA_MQA_LOGITS_CHUNK_MAX_ROWS is set, admission estimates the post-chunk peak using that row cap; otherwise it remains conservative and assumes the full extend-row count. Constraint: Scheduler admission must stay CPU-only and cannot query CUDA free memory. Rejected: Add a separate scheduler limit for MQA logits | the existing max-buffer-size knob is the right aggregate admission budget. Rejected: Use SGLANG_NSA_MQA_LOGITS_FREE_MEM_FRACTION in scheduler | that depends on runtime CUDA free memory and would make admission host-sync or stale. Confidence: medium Scope-risk: moderate Directive: Keep the estimator conservative when chunk max rows is unset; do not rely on CUDA free-memory queries in scheduler admission. Tested: Local py_compile for estimator, scheduler, schedule_policy, and estimator tests. Tested: Local pytest test_cp_shared_kv_prefill_buffer_estimator.py: 5 passed. Tested: Remote g0034 cjy-glm5-new py_compile and estimator pytest: 5 passed. Not-tested: ETE scheduler admission under high-cache-hit bs>1 traffic. Co-authored-by: OmX <omx@oh-my-codex.dev> |
||
|
|
ddc1233955 |
Bound CP MQA logits buffers with row chunking
CP shared-KV bs>1 can build large fp32 MQA-logits temporaries from DeepGEMM fp8_mqa_logits. The official SGLang path already chunks normal NSA MQA logits by query rows behind a cached memory budget; carry the same budget control into our NSA indexer and extend it to CP-ragged topk paths that use row-wise topk_indices_offset_override. This keeps the previous one-time cached memory-budget behavior rather than the recent current-free-mem per-forward variant that regressed performance. A new optional max-rows env provides an explicit hard cap for debugging or controlled ETE runs without adding host syncs. Constraint: DeepGEMM materializes fp32 [q, k] logits internally, so row chunking is the narrowest way to cap temporary memory Rejected: Restore the reverted syh current-free-mem implementation | it changed hot-path heuristics and showed poor runtime performance Rejected: Split by K/context dimension | would change topk semantics and require a different transform contract Confidence: medium Scope-risk: moderate Directive: CP-ragged chunking relies on topk_indices_offset_override being row-addressed; do not route non-ragged CP paths through it without separate validation Tested: Local py_compile for environ.py, nsa_indexer.py, and test_cp_shared_kv_runtime.py Tested: Remote g0034 cjy-glm5-new py_compile for environ.py, nsa_indexer.py, and test_cp_shared_kv_runtime.py Tested: Remote pytest TestCpSharedKVTaiMaterializeIntegration, 17 passed Not-tested: CUDA ETE high-cache-hit bs>1 workload memory/performance after chunking Co-authored-by: OmX <omx@oh-my-codex.dev> |
||
|
|
cc908dd556 |
Reuse DeepGEMM JIT cache across restarts
New sgl-deep-gemm wheels persist cubins, but SGLang was still replaying the full warmup envelope after every restart. The wrapper also translated SGLANG_DG_* env after importing deep_gemm, so the wheel could observe stale cache settings.\n\nMove DeepGEMM JIT env preparation before the first import, bridge the current SGLANG_DG_USE_NVRTC spelling, replace dense M walks with category-specific sparse M grids, and add a SGLang-side warmup manifest under the DeepGEMM cache dir. A manifest hit skips only the SGLang replay loop; missing or stale cache entries still force warmup and refresh the manifest.\n\nConstraint: sgl-deep-gemm 0.1.2 removed the compile-mode API, so dense 1..m_max loops launch real kernels.\nConstraint: DeepGEMM cache keys include compiler flags such as the deep_gemm include path; path-stable environments are still required for cross-container cache hits.\nRejected: SGLANG_JIT_DEEPGEMM_PRECOMPILE=0 | bypasses warmup instead of making cache reuse correct.\nRejected: Enable per-process cache dirs by default | isolates corruption but defeats persistent cross-process cache reuse.\nConfidence: medium\nScope-risk: moderate\nDirective: Do not reintroduce dense DeepGEMM warmup for new wheels unless compile-mode support is verified on the target wheel.\nTested: Local py_compile for jit_cache.py configurer.py compile_utils.py model_runner.py\nTested: Local pytest test/registered/unit/layers/test_deep_gemm_jit_cache.py (6 passed)\nTested: Remote g0034 cjy-glm5-new py_compile for same files\nTested: Remote g0034 cjy-glm5-new pytest test/registered/unit/layers/test_deep_gemm_jit_cache.py (6 passed)\nTested: Remote env probe verified SGLANG_DG_CACHE_DIR and SGLANG_DG_USE_NVRTC bridge to DG_JIT_* before configurer completes\nNot-tested: Full decode cold-cache/hot-cache restart timing and manifest-hit ETE log validation |
||
|
|
3a43727216 |
Bound CP prefill batching by estimated temp memory
CP shared-KV bs>1 batching was only bounded by request count, extend tokens, and cached tokens. That left temporary GPU buffers such as MLA/index materialization, remap metadata, logits windows, and transfer descriptors implicit, and raw extend-token limits could exceed the active chunked-prefill budget.\n\nThis adds an explicit max-buffer-size admission gate with a CPU-only stream-aware estimator, wires it through PrefillAdder/Scheduler, performs a startup CUDA smoke allocation when configured, and reports the estimate in the scheduler admission benchmark. When chunked prefill is active, the effective CP extend-token limit is capped by the current chunk budget so the CP path does not advertise unreachable batch capacity or lift max-prefill-tokens too far.\n\nConstraint: Admission estimation must stay CPU-only on the scheduler hot path; CUDA allocation is limited to startup smoke checking.\nConstraint: Single oversized requests must still be allowed to run alone to avoid scheduler deadlock.\nRejected: Rely only on --max-prefill-tokens | it does not reliably bound the first oversized request and does not model cache-hit/load-back pressure.\nRejected: Let CP extend limit exceed chunked-prefill size | it creates an unreachable effective capacity and misleading budget lift.\nConfidence: medium\nScope-risk: moderate\nDirective: If bs>1 L1 prefetch is enabled later, update CPSharedKVPrefillBufferEstimatorContext.bs_gt1_l1_prefetch_enabled and include the live prefetch dense buffers in overlap windows.\nTested: local py_compile for touched files\nTested: local PYTHONPATH=python pytest -q test/registered/unit/managers/test_cp_shared_kv_prefill_buffer_estimator.py (4 passed)\nTested: remote cjy-glm5-new targeted pytest for new server_args, PrefillAdder, estimator, and benchmark cases (10 passed)\nTested: remote cjy-glm5-new PYTHONPATH=python pytest -q test/registered/unit/managers/test_cp_shared_kv_prefill_buffer_estimator.py test/registered/unit/managers/test_prefill_adder.py test/registered/unit/managers/test_prefill_scheduler_admission_bench.py (29 passed before chunk cap, then test_prefill_adder.py 21 passed after chunk cap)\nNot-tested: full server_args suite because existing TestPrepareServerArgs tries to reach HuggingFace and fails under container DNS/network\nNot-tested: GLM5 ETE smoke with --cp-shared-kv-prefill-max-buffer-size |