Stabilize CP HiCache page-tail ownership under EAGLE reuse
CP shared KV and HiCache now keep page-aligned physical ownership while preserving valid-token radix semantics. Repeated tiny EAGLE exact hits free duplicate tail pages instead of leaking one allocator page, owner-lane load-back uses page-vector admission/eviction, and single-DP idle schedulers avoid entering an unnecessary MLP-sync collective. The commit also records the current page-aligned cache contract and adds gated decode-side EAGLE accept diagnostics so future accept-length collapses can be tied to draft KV/state transfer evidence instead of more prefill cache speculation. Constraint: CP HiCache allocator ownership is page-granular while radix matching remains valid-token based. Constraint: New diagnostics must be gated and must not alter normal EAGLE, transfer, or cache behavior. Rejected: Padding short requests to cp_size or 2*cp_size pages | wastes KV capacity and still hides valid-tail lifecycle bugs. Rejected: Adding more unconditional collectives to prove CP consistency | hot-path collectives previously caused severe performance risk. Confidence: medium Scope-risk: broad Directive: Do not reintroduce silent fallback for CP shared KV/HiCache paths; warning-level fallback or fail-fast is intentional. Tested: git diff --check Tested: local py_compile for all modified Python files Tested: remote g0034 container py_compile for modified Python/test files Tested: remote g0034 container PYTHONPATH=python python -m pytest -q test/registered/unit/layers/test_nsa_cp_utils.py test/registered/unit/mem_cache/test_cp_shared_kv_runtime.py test/registered/unit/mem_cache/test_cp_hicache_load_back_owner_lanes.py test/registered/unit/managers/test_scheduler_dp_attn_mixin.py => 114 passed, 5 warnings, 2 subtests passed Not-tested: full ETE traffic rerun after this commit Not-tested: CUDA/TAI kernel benchmark coverage for all production shapes
This commit is contained in:
@@ -509,6 +509,22 @@ class DecodePreallocQueue:
|
||||
len(kv_args.state_data_ptrs),
|
||||
)
|
||||
|
||||
if envs.SGLANG_EAGLE_ACCEPT_DEBUG.get() and self.spec_algorithm.is_eagle():
|
||||
logger.info(
|
||||
"[EAGLE_ACCEPT_DEBUG] decode_kv_manager cp_rank=%s "
|
||||
"target_kv_bufs=%s draft_kv_bufs=%s total_kv_bufs=%s "
|
||||
"target_state_type=%s registered_state_bufs=%s "
|
||||
"draft_state_type=%s draft_state_bufs=%s",
|
||||
self.tp_rank,
|
||||
target_kv_buffer_count,
|
||||
draft_kv_buffer_count,
|
||||
len(kv_args.kv_data_ptrs),
|
||||
kv_args.state_type,
|
||||
len(kv_args.state_data_ptrs),
|
||||
kv_args.draft_state_type,
|
||||
kv_args.draft_state_buffer_count,
|
||||
)
|
||||
|
||||
kv_args.ib_device = self.scheduler.server_args.disaggregation_ib_device
|
||||
kv_args.gpu_id = self.scheduler.gpu_id
|
||||
kv_manager_class = get_kv_class(self.transfer_backend, KVClassType.MANAGER)
|
||||
|
||||
@@ -350,6 +350,22 @@ class PrefillBootstrapQueue:
|
||||
len(kv_args.state_data_ptrs),
|
||||
)
|
||||
|
||||
if envs.SGLANG_EAGLE_ACCEPT_DEBUG.get() and self.spec_algorithm.is_eagle():
|
||||
logger.info(
|
||||
"[EAGLE_ACCEPT_DEBUG] prefill_kv_manager cp_rank=%s "
|
||||
"target_kv_bufs=%s draft_kv_bufs=%s total_kv_bufs=%s "
|
||||
"target_state_type=%s registered_state_bufs=%s "
|
||||
"draft_state_type=%s draft_state_bufs=%s",
|
||||
self.tp_rank,
|
||||
target_kv_buffer_count,
|
||||
draft_kv_buffer_count,
|
||||
len(kv_args.kv_data_ptrs),
|
||||
kv_args.state_type,
|
||||
len(kv_args.state_data_ptrs),
|
||||
kv_args.draft_state_type,
|
||||
kv_args.draft_state_buffer_count,
|
||||
)
|
||||
|
||||
kv_manager_class = get_kv_class(self.transfer_backend, KVClassType.MANAGER)
|
||||
kv_manager = kv_manager_class(
|
||||
kv_args,
|
||||
|
||||
Reference in New Issue
Block a user