Stabilize CP HiCache page-tail ownership under EAGLE reuse

CP shared KV and HiCache now keep page-aligned physical ownership while preserving valid-token radix semantics. Repeated tiny EAGLE exact hits free duplicate tail pages instead of leaking one allocator page, owner-lane load-back uses page-vector admission/eviction, and single-DP idle schedulers avoid entering an unnecessary MLP-sync collective.

The commit also records the current page-aligned cache contract and adds gated decode-side EAGLE accept diagnostics so future accept-length collapses can be tied to draft KV/state transfer evidence instead of more prefill cache speculation.

Constraint: CP HiCache allocator ownership is page-granular while radix matching remains valid-token based.

Constraint: New diagnostics must be gated and must not alter normal EAGLE, transfer, or cache behavior.

Rejected: Padding short requests to cp_size or 2*cp_size pages | wastes KV capacity and still hides valid-tail lifecycle bugs.

Rejected: Adding more unconditional collectives to prove CP consistency | hot-path collectives previously caused severe performance risk.

Confidence: medium

Scope-risk: broad

Directive: Do not reintroduce silent fallback for CP shared KV/HiCache paths; warning-level fallback or fail-fast is intentional.

Tested: git diff --check

Tested: local py_compile for all modified Python files

Tested: remote g0034 container py_compile for modified Python/test files

Tested: remote g0034 container PYTHONPATH=python python -m pytest -q test/registered/unit/layers/test_nsa_cp_utils.py test/registered/unit/mem_cache/test_cp_shared_kv_runtime.py test/registered/unit/mem_cache/test_cp_hicache_load_back_owner_lanes.py test/registered/unit/managers/test_scheduler_dp_attn_mixin.py => 114 passed, 5 warnings, 2 subtests passed

Not-tested: full ETE traffic rerun after this commit

Not-tested: CUDA/TAI kernel benchmark coverage for all production shapes
This commit is contained in:
laoyao0822
2026-05-30 01:20:01 +08:00
parent 21065cdfdf
commit b56a4f2e6b
16 changed files with 964 additions and 17 deletions
@@ -509,6 +509,22 @@ class DecodePreallocQueue:
len(kv_args.state_data_ptrs),
)
if envs.SGLANG_EAGLE_ACCEPT_DEBUG.get() and self.spec_algorithm.is_eagle():
logger.info(
"[EAGLE_ACCEPT_DEBUG] decode_kv_manager cp_rank=%s "
"target_kv_bufs=%s draft_kv_bufs=%s total_kv_bufs=%s "
"target_state_type=%s registered_state_bufs=%s "
"draft_state_type=%s draft_state_bufs=%s",
self.tp_rank,
target_kv_buffer_count,
draft_kv_buffer_count,
len(kv_args.kv_data_ptrs),
kv_args.state_type,
len(kv_args.state_data_ptrs),
kv_args.draft_state_type,
kv_args.draft_state_buffer_count,
)
kv_args.ib_device = self.scheduler.server_args.disaggregation_ib_device
kv_args.gpu_id = self.scheduler.gpu_id
kv_manager_class = get_kv_class(self.transfer_backend, KVClassType.MANAGER)
@@ -350,6 +350,22 @@ class PrefillBootstrapQueue:
len(kv_args.state_data_ptrs),
)
if envs.SGLANG_EAGLE_ACCEPT_DEBUG.get() and self.spec_algorithm.is_eagle():
logger.info(
"[EAGLE_ACCEPT_DEBUG] prefill_kv_manager cp_rank=%s "
"target_kv_bufs=%s draft_kv_bufs=%s total_kv_bufs=%s "
"target_state_type=%s registered_state_bufs=%s "
"draft_state_type=%s draft_state_bufs=%s",
self.tp_rank,
target_kv_buffer_count,
draft_kv_buffer_count,
len(kv_args.kv_data_ptrs),
kv_args.state_type,
len(kv_args.state_data_ptrs),
kv_args.draft_state_type,
kv_args.draft_state_buffer_count,
)
kv_manager_class = get_kv_class(self.transfer_backend, KVClassType.MANAGER)
kv_manager = kv_manager_class(
kv_args,