Commit Graph
62 Commits
Author SHA1 Message Date
leaveletandClaude Fable 5 47cb9b9d11 Make the duplicate-reservation ack gate O(1) for fresh node ids
node_has_undrained_write_ack scanned ack_write_queue on every
reservation, including the per-request-per-chunk prepare path whose
node ids are freshly minted and can never be queued. Track the max
node id ever appended to the queue: any queued id is <= max by
construction (no monotonicity assumption needed for correctness), so
fresh ids exit on one integer compare and the scan remains only for
re-reservations of old ids (the rare write_backup fallback).

All three ack_write_queue append sites now go through a single
_append_write_ack funnel that maintains the max — the zero-owned-rank
and non-CP write acks previously would have bypassed it, allowing
false negatives on ranks owning no pages of a node.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-11 10:35:29 +00:00
leaveletandClaude Fable 5 698a3e2431 Enforce one in-flight CP write ack per node at the producer
dfa168abe9 made duplicate ack completion idempotent at the consumer
(writing_check), which is correct and lock-balance-safe but masks the
producer invariant breach: two acks can only coexist for one radix
registration when an ack is orphaned in ack_write_queue after a
rollback cleared ongoing_write_through/pending_host_backups,
re-opening the _node_host_write_pending guard for a fresh
registration. _rollback_pending_backup (write_backup's exception
path) was the one rollback that did not scrub the node's acks — and
it also left a half-submitted layer-write state alive, which later
forwards would keep driving, writing D2H into host slots the rollback
had already evicted.

Close the invariant structurally:
- reserve_write_cp refuses (HiCacheWriteFailure, the existing
  skip-this-round path both callers already handle) when the node
  still has an undrained final ack, computed by scanning the small
  ack queue — no new state to keep in sync.
- _rollback_pending_backup now cancels the pending layer-write state
  (write-stream sync so in-flight per-layer copies finish before the
  host slots are evicted) and scrubs the node's queued acks via the
  scrub factored out of _rollback_prepared_cp_backup.
- The consumer-side duplicate guard from dfa168abe9 is kept as
  defense in depth.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-11 10:22:26 +00:00
leaveletandClaude Fable 5 75d7d8772e Unify over-length errors into the PayloadTooLargeError 413 format
Over-long inputs produced two different client errors depending on
which bound rejected them: the TokenizerManager pre-check (raw
context_len) returned 413 PayloadTooLargeError ('The input (N tokens)
is longer than the model's context length (M tokens).'), while inputs
between that and the scheduler's stricter effective limit hit
validate_input_length and returned 400 BAD_REQUEST with different
wording (and a confusing 'X exceeds X' message since the check is >=).

Unify on the 413 format end to end:
- validate_input_length wording now matches the TokenizerManager
  message, reporting the effective per-request limit.
- set_finish_with_abort takes status_code/err_type; the scheduler
  length-rejection sites abort with REQUEST_ENTITY_TOO_LARGE +
  PayloadTooLargeError. The batch handler previously queued the
  over-long request WITHOUT marking it aborted (it proceeded to
  prefill) — also fixed.
- Non-streaming aborts with 413 raise PayloadTooLargeError (now a
  ValueError subclass so raw /generate-style endpoints that only
  catch ValueError still respond; the OpenAI layer's except clause
  is reordered to win and emit the 413 format).
- Streaming abort responses prefer the scheduler-provided err_type
  over the HTTPStatus name.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-11 08:01:02 +00:00
leaveletandClaude Fable 5 d01601c171 Demote per-chunk CP per-layer transfer register log to debug
'[CP_PER_LAYER_TRANSFER] registered room=... chunk=...' fired at INFO
once per room per chunk on every prefill send. The one-time manager
registration stays INFO and failure paths stay WARNING; the
success-side finish log was already debug.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-11 06:25:37 +00:00
leaveletandClaude Fable 5 afdd1d0992 Demote per-request load-back log to debug
init_load_back logged 'loading back N tokens for node M' at INFO for
every request triggering an L2->L1 HiCache load-back (hi_mamba already
had it at debug). Demote and switch both to lazy %-formatting so the
f-string is not built when debug is off.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 10:22:35 +00:00
leaveletandClaude Fable 5 3eef989739 Harden IndexCache for first e2e: validation, fail-fast, prefetch gating
Pre-e2e audit fixes for NSA index-topk sharing (index_topk_pattern):

- Validate the pattern at init: F/S charset only, must start with F
  (layer 0 has nothing to share from), and length must equal
  num_hidden_layers — a short pattern previously crashed deep in pool
  init with an obscure per-layer ValueError, a long one silently
  ignored tail characters. Warn when a pattern shadows a configured
  index_topk_freq (pattern takes precedence silently otherwise).
- Fail fast when a shared (S) layer receives prev_topk_indices=None:
  both gated indexer call sites previously fell through to omitting
  topk_indices entirely, running sparse attention without indices and
  silently corrupting output if threading were ever dropped.
- Reject pipeline parallelism with index-topk sharing (same rationale
  as the existing TBO guard): topk indices are not threaded across
  PP-stage boundaries, and each stage's loop resets them to None.
- Gate the CP shared-KV index prefetcher on the next layer having an
  index-cache slot: prefetching an S layer's nonexistent buffer
  tripped the inactive-layer fail-fast and disabled the prefetcher
  for the entire run on the first F->S transition.
- Log the resolved active index layer plan at pool init so an e2e run
  can confirm IndexCache is actually on.

Tests: pattern validation cases incl. the GLM-5 reference 78-char
pattern, next_skip == skip-of-next invariant; existing 'C' pattern
test updated to upstream 'F' charset.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 06:17:52 +00:00
leaveletandClaude Fable 5 5bec69f40a Drop redundant np.isin from CP-rank KV page filter
filter_kv_indices_for_cp_rank built rank_page_indices =
kv_indices[range_mask] and then ran np.isin(kv_indices,
rank_page_indices) — provably identical to range_mask itself (a value
is in the filtered subset iff it passes the same range test), at the
cost of an extra sort per chunk send under
SGLANG_DISAGGREGATION_ALL_CP_RANKS_TRANSFER=1. Apply the range mask
directly: 30.5 -> 8.1 us per 1024-page chunk. Differential test pins
equivalence against a verbatim copy of the old logic.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 05:15:16 +00:00
leaveletandClaude Fable 5 b742371fe3 Skip get_load() snapshot when no DP load-balancer consumes it
stream_output_generation computed a full load snapshot on every
output flush: get_load() walks the waiting/bootstrap/prealloc queues
summing seqlens and builds a DpRequestInfoqOutput per queued request.
The only consumer is the DP load-balancer path in tokenizer_manager,
which forwards it iff dp_size > 1. Gate the computation on the same
condition; at dp_size == 1 the snapshot was computed and shipped for
nothing on every flush.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 05:12:40 +00:00
leaveletandClaude Fable 5 e6864fbea7 Demote HiCache/CacheCtrl per-op logs from info to debug
The [HiCache-write]/[HiCache-load]/[HiCache-collective] and
[CacheCtrl-write] per-operation logs fired at INFO on the scheduler
thread during serving (every write-through, load-back, ack release,
and rate-limited collective summary), eating into stream gaps.
Demote all 29 to debug; one-time [HiCache-draft] init logs stay INFO
and genuine failures stay WARNING (write_backup CP FAILED after
deterministic retry).

Also repair TestHiCacheEvictLoggingLevels, which asserted markers
that no longer exist in this tree (stale list, failing on the
pristine branch): pin the current hot-path markers as debug-only and
extend the check to cache_controller.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 05:10:45 +00:00
leaveletandClaude Fable 5 d221212565 Batch EAGLE spec output D2H copies in disagg prefill results
process_batch_result_disagg_prefill did one synchronous
hidden_states[i].cpu().clone() per finished request, and stored
topk_p/topk_index as GPU row views whose D2H then happened one by one
inside MetadataBuffers.set_buf — 3*bs device syncs per finished batch
on the scheduler thread. Gather the finishing rows once and do one
D2H per tensor (3 total); requests hold CPU row views, which set_buf
copies CPU-to-CPU into the metadata buffers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 05:07:51 +00:00
leaveletandClaude Fable 5 87a22b17ce Fuse SamplingBatchInfo tensor construction into one pass
from_schedule_batch built temperature/top_p/top_k/min_p (+seed) with
4-5 separate list comprehensions and one synchronous H2D copy each,
plus 4 more passes for the is_all_greedy/need_* flags. Collect
everything in a single pass over reqs and upload the float params as
one pinned non-blocking H2D copy (disjoint device views of one
buffer; filter/merge only index and cat, producing fresh tensors, so
the shared buffer is safe), int32 top_k and optional int64 seeds as
their own pinned copies.

B300 (torch 2.11 cu130), scheduler-thread blocking time per call:
  bs=8: 38.8 -> 18.2 us (2.1x); bs=32: 1.9x; bs=200: 1.2x
CPU-only construction at bs=200: 2890 -> 1387 us (2.1x).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 05:04:39 +00:00
leaveletandClaude Fable 5 0d065a8ab0 Speed up Python bigram key conversion with C-level zip
The EAGLE bigram fallback built 64K-token keys with an index-based
Python comprehension (6.3 ms per call at 64K tokens). zip + islice
runs the pairing loop in C and avoids copying the shifted operand:
3.6-4.2 ms per call (~1.6x). Output is byte-identical (same tuples);
the tai-kernel fast path is unaffected.

Packing bigrams into int64s via numpy was measured and rejected: the
list<->array boxing makes it slower (4.1 ms) than zip until token ids
are numpy end-to-end, and it would change HiCache storage hash inputs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 04:56:24 +00:00
leaveletandClaude Fable 5 8979c81a22 Make RadixKey slicing zero-copy to fix quadratic match_prefix
RadixKey.__getitem__ copied the token list on every slice, and the
match/insert tree walks re-slice the remaining key at every node hop,
making a single match O(len * hops) — quadratic for long cached
prefixes. Slices now return O(1) offset-based views over a shared
backing list; the key-match and child-key functions index the backing
list directly so views are never materialized on the hot path.

Keys stored in tree nodes are compacted at every store site (same cost
as the old copying slices), so lock-ref walks, eviction, splits, and
controller-thread reads never observe a key pinning a transient
backing list. Node-key compactness is enforced by a tree-walk test.

Microbenchmark (64K-token full hit, page_size=64):
  16-node path: 1.86 -> 0.87 ms/match (2.1x)
  128-node path: 9.80 -> 1.19 ms/match (8.2x), insert re-walk 6.9x

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 04:28:03 +00:00
leavelet b35495a28d chore(work): journal bs>1 contamination fix + B300 P/D validation
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit c98cc2736d7471c31ee38b22eb5c8392ec6f7519)
2026-06-09 00:01:40 +08:00
leavelet 165b6a01f1 fix(nsa): correct bs>1 CP shared-KV KV per-request page_inverse + CP/EAGLE compute-padding
Two correctness fixes in the bs>1 CP shared-KV prefill path, validated end-to-end on B300 (GLM-5.1-FP8, EAGLE, fp8 KV, CP8):

1) Cross-request KV contamination: the value-keyed 1-D page_inverse was last-writer-wins, so two requests sharing a physical page (radix shared-prefix / current-reuse) aliased each other's KV. Replace with a per-request flat page_inverse[batch_rows*capacity] indexed req_id*capacity+lp, threaded through build/remap/fill_current/prefetch and the TAI/Triton kernels. req_id sourced via slot//pages_per_request (build), repeat_interleave (logical_locs), and a companion tensor through the same CP valid-split (current_locs).

2) CP+EAGLE compute-padding cache-write: forward_absorb_prepare's rebuild_cp_kv_cache all-gathers current k back to global rows under current-reuse, but the compute-padding branch fed that full k straight to select_cp_local_valid_rows_for_cache_write (which requires per-rank compute rows) -> fail-fast. Add cp_localize_current_kv_to_compute_rows to re-localize via the batch-plan compute split first (mirrors the non-padding branch).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 28efef91b05ede1a7072ee451e2aea39ecc3a5bd)
2026-06-09 00:01:04 +08:00
leavelet 81a0191331 perf(deepgemm): CP-aware dense warmup M-grid for NSA prefill CP
Shrink the non-grouped (dense/attention) DeepGEMM warmup M grid by attn_cp_size under NSA prefill in-seq CP (per-rank M = tokens/cp), while keeping the grouped MoE GEMM grid full (deepep all-to-all re-gathers all tokens; topk==ep_size keeps MoE M ~= chunked). Gated by _cp_dense_warmup_divisor.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 672ef3a6609abd100e5e95b42016d17d0a2966e5)
2026-06-08 23:46:59 +08:00
leavelet 134b88b9b4 fix(server_args): keep NSA prefill/decode backends in the same family for fp8
calculate_mla_kv_cache_dim only packs fp8 (override dim -> nsa_kv_cache_store_fp8=True, required by the flashmla_sparse dequant path) when NEITHER prefill nor decode backend is trtllm. The Blackwell-no-DP guard defaulted the unset side to trtllm, producing a mixed flashmla/trtllm pair that left fp8 KV unpacked and crashed flashmla_sparse ('kv must have dtype bf16'). Now an explicit flashmla choice on either side pulls the unset side to flashmla too.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit ed30312c3d599595e40ae99bc2832159b634cca6)
2026-06-08 23:46:59 +08:00
leavelet 98fdae3375 fix(config): restore GlmMoeDsa index_topk_freq dropped by HF config load
GlmMoeDsaConfig.__init__ has no **kwargs, so transformers drops the raw
config.json fields the IndexCache/DSA path reads via getattr (index_topk_freq,
index_topk_pattern, index_skip_topk_offset) and may clobber qk_rope_head_dim —
only named fields like index_topk survive. Re-read them from config.json and
restore in get_config for GlmMoeDsaForCausalLM, matching upstream #27114.

Without this, setting index_topk_freq in the model config has NO effect
(getattr falls back to 1 -> IndexCache stays off). Verified on the B300 image:
get_config("/ssd/models/GLM-5.1-FP8") now returns index_topk_freq=4.

Refs: WI-2026-06-07-001

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 92ff8b47470c61462e252bc7ba428203ff3c324c)
2026-06-08 23:46:59 +08:00
leavelet ce3b5bc874 feat(nsa): port IndexCache (indexer topk sharing across layers)
Port upstream IndexCache for DeepSeek-V3.2 / GLM-5 (#21405 + #27114 gate fix) to
our diverged tree, matching upstream HEAD's merged form. The NSA indexer topk is
computed every `index_topk_freq` layers; `skip_topk` layers reuse the previous
layer's topk via prev_topk_indices threaded through the model forward. Default
index_topk_freq=1 => no sharing => zero behavior change until the model config opts in.

- forward_mla.py: gate the two indexer call sites
  (`if not skip_topk or (is_nextn and prev_topk_indices is None)`), else reuse
  prev_topk_indices; forward_absorb_core returns (output, topk_indices) when
  next_skip_topk is set (topk_indices already threaded prepare->core via inner_state).
- deepseek_v2.py: AttentionMLA.__init__ skip_topk/next_skip_topk setup
  (freq/pattern/offset; is_nextn=True/True); thread prev_topk_indices through
  forward/forward_prepare/op_core; DecoderLayer returns a 3-tuple (tuple-unpack
  placed AFTER the CP shared-KV finally); Model.forward loop threads topk_indices.
- deepseek_nextn.py: 3-tuple decoder unpack.
- server_args.py: port the #27114 guard - raise on --enable-two-batch-overlap with
  index-topk sharing (the TBO op path does not propagate topk across layers, so
  shared layers would run sparse attention with no indices).

All DecoderLayer.forward callers covered (Model loop, nextn, glm4_moe_lite /
mistral_large_3_eagle inherit, TBO op-path discards via op_core). Did NOT port
#24392 (orthogonal indexer-topk capture/output infra). Import-validated on the
B300 / torch-2.11 image. Enabling it requires setting `index_topk_freq` in the
model config (Zhipu-confirmed for GLM-5.1).

Refs: WI-2026-06-07-001

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit af2a41b8d396033e378b6bdd5f6baebf2a98c301)
2026-06-08 23:46:59 +08:00
leavelet 4cf8885a74 perf(nsa): cache MQA-logits budget + drop indexer rope clones
Two upstream NSA-indexer perf ports (verified against our diverged tree), cutting
CPU launches / host syncs on the prefill critical path for GLM-5.1-FP8 on B300:

- #25299 (beaff00331): cache the MQA-logits chunking budget per device so prefill
  stops calling torch.cuda.mem_get_info (a host sync) on every indexer pass. The
  budget is computed once and capped by the mem_fraction_static serving headroom;
  a static guard is used (uncached) during cuda-graph capture, and the first real
  prefill caches the free-memory budget.
- #22232 (671fe73961): replace `dst = src.clone()` slice write-backs of the RoPE
  output with a data_ptr-guarded `dst.copy_(src)`. q_rope/k_rope are torch.split
  views of query/key, so when RoPE runs in place src/dst alias and the write-back
  is a redundant no-op (guard skips it); otherwise one copy instead of clone+copy.
  Saves an alloc+copy per q/k per indexer call. (Skipped the PR's AMD-only
  @torch.compile cleanup.)

Both verified to import on the B300 / torch-2.11 image. Deliberately NOT taken:
#21332 (un-force MHA one-shot on Blackwell — risky, per user), #23856 (torch.mm
indexer GEMM — per user).

Refs: WI-2026-06-07-001

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 6bc23945e4ff5c68b0618173856c7daf4653fa10)
2026-06-08 23:46:59 +08:00
leavelet db5ab48b19 feat(deepgemm): migrate to sgl-deep-gemm 0.1.2 wheel API
Port the DeepGEMM "deprecate-from-sgl-kernel, separate sgl-deep-gemm wheel"
migration (upstream PR #24268 and follow-ons) so our code runs against the
torch-2.11 dev-cu13 image, which ships DeepGEMM as the separate sgl-deep-gemm
0.1.2 wheel (import name still `deep_gemm`). Verified against upstream/main HEAD,
NOT the introducing PR — the wheel API drifted between 0.0.1 and 0.1.2.

Ports (all verified against HEAD = wheel 0.1.2):
- nsa_indexer/nsa_backend: paged-MQA context_lens must be (N_total, 1). The
  0.0.1 form (batch_size, next_n) DEADLOCKS fp8_paged_mqa_logits on next_n>=2
  (our EAGLE deploy uses next_n=4 on SM90/H200, which does not take the SM100-
  only DG-native broadcast path). Matches HEAD's _to_2d_context_lens.
- compile_utils warmup: hasattr-guard the dropped get/set_compile_mode API;
  pass m_indices positionally to m_grouped_fp8_gemm_nt_contiguous.
- fp8_utils.transform_scale_ue8m0: restore TMA-aligned stride when the DLPack
  round-trip collapses a size-1 trailing dim.
- moe_runner/deep_gemm: guard the SBO masked-gemm return unpack when overlap is
  inactive (#26839) — reachable via our --enable-single-batch-overlap.
- entrypoint: enable DeepGEMM PDL by default, hasattr-guarded (#23979).
- engine: require sglang-kernel >= 0.4.3 (first sgl-deep-gemm-era kernel); the
  migrated (N_total,1) paths would misbehave on the old bundled DeepGEMM.

Verified-unchanged symbols left as-is: fp8_mqa_logits, fp8_gemm_nt, bf16_gemm_*,
get_mk_alignment_for_contiguous_layout, transform_sf_into_required_layout,
get/set_num_sms, the masked-gemm signature. Runtime validation pending the
torch-2.11 image (tai-kernel rebuild + harness).

Refs: WI-2026-06-07-001

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 32818784d332b5c48faf8027247cd4cd9a0f48cd)
2026-06-08 23:46:59 +08:00
leaveletandClaude Opus 4.8 a1d5652c7a chore(work): record stub-shadowing fix + nsa_pool_host env-failure findings
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 10:16:01 +00:00
leaveletandClaude Opus 4.8 b1cbacffae test(cp_shared_kv): load real sgl_kernel first to stop CPU-CI stub leaking
test_cp_shared_kv_runtime installs CPU-CI sgl_kernel stubs at import time via
sys.modules.setdefault + torch.library FRAGMENT defs. On a GPU box, when this
module is collected before the real sgl_kernel loads, the empty stub permanently
shadows it process-wide, so a later real-kernel test (e.g.
fast_topk_transform_ragged_fused in test_nsa_topk_transform) calls a None-returning
lambda and crashes. Importing the real sgl_kernel first makes setdefault keep the
real module (setattr only fills missing names) and the FRAGMENT defs hit the
already-registered path; on CPU-CI the import fails and stubbing proceeds as before.

Verified on GPU: test_cp_shared_kv_runtime alone 120 passed; combined with
test_nsa_topk_transform 125 passed (previously 1 failed / segfault on cleanup).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 10:15:35 +00:00
leaveletandClaude Opus 4.8 899828fe22 chore(work): record per-layer chunked fix, guard, cache-precision debug, rebase
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:53:45 +00:00
leaveletandClaude Opus 4.8 24bafabea3 chore: gitignore docs_internal and tai-kernel (internal harness/notes/kernels)
These hold the internal benchmarking harness, investigation notes, and local
kernel sources used during development. They were inadvertently committed in
earlier lever-A work and have been stripped from history; ignore them so they
stay on disk for local use but never get tracked again.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:52:39 +00:00
leaveletandClaude Opus 4.8 243f1b964c fix(cp_per_layer_transfer): gate per-layer path on a single non-dummy decode info
The transfer worker iterates every non-dummy decode info for a room and calls
per_layer_mgr.finish() once per info, but register_per_layer_transfer registers
exactly one context per room/chunk (built for one info's dst_kv_indices). This is
only sound when there is exactly one non-dummy info (required_dst_info_num == 1).
With decode attn_tp < prefill attn_tp a single prefill rank holds >1 non-dummy
infos; finishing once-per-info would over-pop chunk contexts and under-deliver KV
to the other infos. Make the assumption explicit: register only when there is one
non-dummy info, otherwise fall back to the monolithic post-forward transfer (which
fans out to all infos correctly). Found by an independent first-principles audit.

Adds TestRegisterGuardSingleInfo (2-info fallback, 1-info register, all-dummy
fallback) exercising the real MooncakeKVManager.register_per_layer_transfer.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 e9354c41bc fix(cp_per_layer_transfer): use absolute chunk_page_start for chunked KV dst mapping
The per-layer KV transfer registration hardcoded chunk_page_start=0 when
filtering CP-shared-KV owned pages. The CP filter's second return (`positions`)
are absolute full-sequence page positions built from chunk_page_start, and the
transfer indexes the FULL-request dst_kv_indices by those absolute positions
(mirroring the monolithic send(), which passes chunk_page_start=index_slice.start
— the cumulative page offset). With start=0, chunk N>0's positions were
chunk-local, so its KV was written onto chunk 0's decode pages, corrupting the
decode output. Non-chunked requests (single chunk, start=0) were unaffected,
matching the observed symptom (non-chunked byte-identical, chunked garbage).

Fix: chunk_page_start = chunk_key // page_size, where chunk_key is the chunk's
start_send_idx (page-aligned), making it exactly the monolithic index_slice.start.

Verified: opus first-principles code audit; empirical mapping-invariant on the
deployed modules (per-chunk == whole-request for all 8 CP ranks; old start=0
sends chunk1 to chunk0's dst); 2 new regression tests (TestChunkedDstMapping).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 beffd3a82f feat(disagg): unify per-layer transfer over chunked prefill (lever A)
Per code review (HiCache load is layer-by-layer & correctly ordered into the compute
stream before the backup hook; current-reuse is a within-forward read that doesn't
rewrite pool pages): unify registration to the EXACT range this forward's
send_kv_chunk transmits — req_to_token[start_send_idx:end_idx], page-floored for a
non-last chunk. Non-chunked = one full range; chunked = one range per chunk. Drop
the is_chunked/start_send_idx skip.

To avoid the review's collision risk (chunk N still finishing when chunk N+1
registers), the manager keys contexts per (room, start_send_idx): _active[room] is a
FIFO deque of (chunk_key, ctx); register dedups the same chunk but appends a new one;
finish(room) pops the FRONT (chunks finish in send order — no chunk key needed in the
transfer_worker); drop drains all the room's chunks; on_layer_end enqueues for all
active chunk contexts (per-ctx note_enqueued dedup keeps each chunk's own events).

28 unit tests pass incl. chunked FIFO + per-chunk dedup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 e12afe8ced perf(disagg): coarse-grained per-layer transfer + skip re-registration (lever A)
For large bs (production target: bs~10 x ~100k tokens), per-layer granularity does
10x79 = 790 submitTransfer calls + CUDA events + enqueues per forward on the forward
thread. Two overhead cuts:
- Group SGLANG_CP_SHARED_KV_PER_LAYER_GROUP (default 8) consecutive layers into ONE
  RDMA submit: ~num_layers/K submits + events + enqueues instead of per-layer; same
  bytes (page index lists are identical across layers). on_layer_end is O(1) at
  non-boundary layers. The last partial group enqueues via the num_layers boundary;
  any misses fall back to one batched sync submit.
- Scheduler hook skips reqs already registered (bs>1 batch-forming re-iterates the
  same reqs ~9x -> was rebuilding the CP filter + context every time).

27 unit tests pass incl. grouping-boundary + batched-fallback.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 5fdf439c7e fix(disagg): dedup per-layer enqueues so high-cache-hit can't hang (lever A)
E2E at high cache-hit + concurrency exposed a vicious cycle: a context stays active
until finish, but the notifier re-fires every layer on every subsequent forward and
note_enqueued counted each, so _enqueued (the finish target) grew by the layer count
each forward (target=3042-5538 observed) faster than 4 workers can drain -> finish
never reaches it -> 30s timeout -> context stays active -> repeat. TTFT p99 = 116s.
Fix: note_enqueued(layer_id) dedups per layer (target caps at the layer count); the
first fire for a layer is from the request's own forward so its event is correct.
Also guard register() against overwriting an active room (was leaking + re-registering,
registered=917 for ~100 reqs).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 5791684864 feat(disagg): broaden lever A to full sequence incl cached prefix (high cache-hit)
The premise holds at realistic context: at ~8k tokens + high cache-hit the KV
transfer is ~30% of TTFT (712ms), sequential after the forward. Lever A was scoped
to prefix==0 so it never activated there. Broaden it to cover the FULL sequence
(cached prefix + new tokens): the per-layer backup hook fires after each layer's
full processing, so the HiCache-loaded prefix KV and forward-written new KV are both
final — transferring all of layer L's pages then overlaps the dominant prefix
transfer with the forward. Add a sonnet-bench verify mode for output-equality.
Gated by the flag; correctness verified next.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 76dc60a1d0 test(disagg): finish() latency breakdown instrumentation (lever A A5 root-cause)
Log per-request finish breakdown (wait_workers / fallback / wait_rdma) to prove
WHY the per-layer transfer stage (~193ms) doesn't shrink vs the monolithic (~172ms)
— i.e. whether the workers fall behind the forward (submits land late, no overlap)
or the RDMA itself is the cost. INFO-gated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 64f767bb42 fix(disagg): cover notifier-missed layers via post-forward fallback (lever A)
E2E diagnostic was precise: "finish TIMEOUT processed=78/79", submit_failed=0 —
the per-layer notifier fires 78x but kv_data_ptrs has 79 layers (the 79th is the
MTP/nextn EAGLE buffer: present in kv_data_ptrs so the monolithic send moves it,
but it doesn't fire the per-layer hook). The old completion required all num_layers,
so it both hung to the timeout AND would silently miss that layer's KV (corruption).

Redesign: gate completion on the ACTUAL enqueued count (note_enqueued), and in
finish() SYNCHRONOUSLY transfer any layers the notifier didn't fire for (KV is fully
written post-forward, no event needed). Net: the per-layer set == kv_data_ptrs,
byte-identical to the monolithic send; robust to any model firing fewer hooks than
KV buffers. The fired layers stay overlapped with the forward.

Unit tests updated (25 pass): fallback transfers the missed layers; submit failures
still report -1.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 246dbddac0 fix(disagg): per-layer completion can't hang + worker CUDA device + diagnostics
E2E (clean run) showed finished ret=-1 at exactly the 30s timeout for every
request: finish() hung because _worker_step swallowed submit_layer exceptions
without counting the layer toward completion (so _processed never reached
num_layers). Fixes:
- submit_layer: try/finally that ALWAYS counts the layer (completion can never
  hang on a per-layer error) and LOGS the actual exception.
- PerLayerTransferManager worker_init: torch.cuda.set_device on each worker thread
  (likely cause — event.synchronize() needs the device set on these fresh threads,
  unlike the transfer_worker where A1's engine call worked).
- finish() logs processed/num_layers on timeout to separate exception-failure from
  notifier-undercount.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 501faa8f5b fix(disagg): tensor-safe prefix check + graceful per-layer register (lever A)
E2E caught it: registered=8 (per-layer activated correctly) then the prefill
crashed with "Boolean value of Tensor with no values is ambiguous" — req.prefix_indices
is a tensor and `prefix or []` evaluated its truthiness. Use a None+len check
(safe for None/list/tensor), and wrap each request's registration in try/except so
a setup error degrades to the monolithic transfer instead of crashing the scheduler.
Also guard start_send_idx with getattr.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 ae18e3adc8 feat(disagg): wire per-layer overlapped transfer end-to-end (lever A, A3-step3)
Complete the lever-A hot-path integration behind SGLANG_CP_SHARED_KV_PER_LAYER_TRANSFER:
- PerLayerTransferContext: add num_layers + a completion event so finish() waits
  until ALL layers are processed before wait_batch_transfers (never races ahead of
  the worker threads and silently drops in-flight layers); times out to FAILURE.
- PerLayerTransferManager: has_room (for the swap) + drop (abort/failure drain so
  outstanding RDMA finishes before pages are reclaimed).
- MooncakeKVManager.register_per_layer_transfer: build + register a context before
  the forward, reusing send()'s exact CP filter (no re-derivation).
- transfer_worker: when a room is per-layer-active, wait those transfers (finish)
  instead of the monolithic send_kvcache -- no double-send; aux/state/completion
  unchanged. The skip path drops the context on abort/failure.
- prefill scheduler: _register_per_layer_transfers(batch) before run_batch, scoped
  (first impl) to single-forward, no-cached-prefix requests (the notifier transfers
  forward-written pages; chunked/cached-prefix are HiCache-loaded -> lever B).

Unit-tested (25 cases). e2e output-equality + TTFT verification next.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 aa6acc9485 feat(disagg/mooncake): build_per_layer_context assembly (lever A, A3-step2)
Add MooncakeKVManager.build_per_layer_context: assembles a PerLayerTransferContext
from the SAME CP-filtered (prefill_kv_indices, dst_kv_indices) the post-forward
transfer uses — so the bytes moved are byte-identical to the monolithic path, and
the CP owner mapping is NOT re-derived (eliminating the #1 correctness risk). It
mirrors the MLA branch of _send_kvcache_generic exactly (get_mla_kv_ptrs_with_pp +
group_concurrent_contiguous + build_layer_blocks, verified set_transfer_blocks-
identical). Returns None for MHA / unregistered decode / empty owned set.

Unit-tested (4 cases): per-layer address correctness + the None guards. Remaining
A3-step3: call this in the send/scheduler flow (register before forward, finish
after, skip the main-KV monolithic send), then output-equality + TTFT verification.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 6bd5d5760c feat(disagg): per-layer block-address computation (lever A, A3-step2 core)
Add build_layer_blocks: the pure per-layer transfer-address computation (src/dst
addrs + lengths for layer L's owned page blocks), the core of the context's
get_blocks closure. Mirrors the mooncake set_transfer_blocks math; the page index
lists are identical across layers, so only the per-layer base ptr + item_len
change. Unit-tested (3 cases incl. the cross-layer invariant). 27 per-layer/async
tests green total.

The remaining A3 step assembles get_blocks from the scheduler's per-request data
(transfer_infos dst indices + decode_kv_args_table dst ptrs + out_cache_loc src +
CP owner filter) before run_batch, and reconciles finish() with send_kv_chunk —
the hot-path integration, to be verified by the bitwise-equivalence harness.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 74e22a0721 feat(disagg): register per-layer transfer manager notifier (lever A, A3-step1)
Wire PerLayerTransferManager into the prefill bootstrap queue: when
SGLANG_CP_SHARED_KV_PER_LAYER_TRANSFER is on, create the manager (real
torch.cuda.Event / current_stream) and register its on_layer_end on the KV pool's
layer_backup_notifiers, exposing it as kv_manager.per_layer_transfer_manager. The
forward's per-layer end hook now reaches the manager. Additive and a no-op until a
request registers a transfer context (A3-step2): on_layer_end returns early when no
contexts are active (verified). Remaining A3 (get_blocks + setup-before-run_batch +
finish reconciled with send_kv_chunk) builds on this.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:05 +00:00
leaveletandClaude Opus 4.8 7ac8067db1 feat(disagg): per-layer transfer manager (lever A, A2 orchestration)
Add PerLayerTransferManager: owns a worker thread-pool + the active per-request
contexts for the current forward batch, registered as a KV-pool notifier.
on_layer_end (forward thread) records ONE CUDA event on the compute stream and
enqueues (ctx, layer, event) per active context; workers do the event-wait +
async submit OFF the forward thread. finish(room) waits the request's transfers.
event_factory/current_stream injected for unit-testability (no CUDA needed).

Unit-tested (5 manager cases, 11 total in test_cp_per_layer_transfer.py):
per-active-ctx enqueue with the event recorded on the stream, worker-step submit +
mark-failed-on-exception, finish pop + idempotency, no-op when idle. The A3
scheduler wiring (notifier registration + setup-before-forward + finish-after +
no-double-send reconcile) is the remaining hot-path step; plan locked in
lever-a-implementation-plan.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:04 +00:00
leaveletandClaude Opus 4.8 e0d47bfc41 feat(disagg): per-layer transfer context for overlap (lever A, A2 core)
Add PerLayerTransferContext: the per-request coordinator for overlapped per-layer
KV transfer. submit_layer(layer_id, event) waits the layer's CUDA write event (on
a background thread, never the compute stream) before async-submitting that
layer's RDMA via the G1 path, so the transfer never reads a layer before its
write kernel finished — the core correctness invariant for the forward overlap.
finish() waits all accumulated batch_ids; idempotent per layer; fails closed.

Unit-tested (test_cp_per_layer_transfer.py, 6 cases): event-wait-before-submit
ordering, idempotency, empty-layer skip, finish-waits-all, submit-failure stop,
wait-failure propagation. The scheduler/notifier wiring (A2-wiring + A3) builds
on this.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:04 +00:00
leaveletandClaude Opus 4.8 a09aec30a3 feat(disagg/mooncake): async per-layer KV transfer primitive (lever A foundation)
Add _transfer_layers_async: submit each layer's transfer non-blocking via the
async API (batch_transfer_async_submit), pipelining layers in the RDMA engine,
then wait for all once (wait_batch_transfers). Gated by the new
SGLANG_CP_SHARED_KV_PER_LAYER_TRANSFER env (default off); replaces the monolithic
all-layers batch_transfer_sync on that path. This is the transfer mechanism for
per-layer overlap (lever A) and removes the per-layer blocking-sync tax measured
in B1a; the forward-overlap hook (G2) builds on it next. Uses the safe async API,
never the OnCuda busy-wait/_exit path.

Unit-tested (test_per_layer_transfer.py, 5 cases): one submit per non-empty layer,
single wait-for-all, empty-layer skip, submit-failure drain + return -1, and
wait-status propagation.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:04 +00:00
leaveletandClaude Opus 4.8 1c6279eeac chore(harness): fix router prometheus port; update work item
Router needs an explicit --prometheus-port under --network=host (collides with
the metrics-enabled servers otherwise). Work item captures the full harness
bring-up: GLM-5.1 PD up e2e, 4 ports validated by smoke, #27372 abort validated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:04 +00:00
leaveletandClaude Opus 4.8 9a91c73e09 feat(disagg/mooncake): add non-blocking async transfer wrapper bindings
Add batch_transfer_async_submit() (non-blocking submit -> batch_id) and
wait_batch_transfers() (block-until-all, graceful) to MooncakeTransferEngine,
exposing the mooncake async API for the upcoming per-layer overlapped transfer.

Deliberately avoids batch_transfer_*_on_cuda: its CUDA host callback busy-waits
on the stream and calls _exit(1) on transfer failure (verified in the mooncake
source transfer_engine_py.cpp), which is unsafe for production (a decode-side
abort mid-transfer would crash the prefill). The async submit + wait pair lets
per-layer submits pipeline in the RDMA engine, then waits once on a background
thread with graceful failure.

Both bindings raise a clear upgrade error if the engine predates the API and
degrade to -1 on transient errors. Mock-engine unit tested (8 cases).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:04 +00:00
leaveletandClaude Opus 4.8 494153da1f docs(disagg): per-layer async transfer investigation, plan, and harness
Capture the design trail for the per-layer/async KV transfer work: verified
problem + infra/risk map, mooncake version/CUDA-13 + async-API findings, the
upstream conn.py pick triage, the master plan, the per-layer/async design, the
g0034 env setup, and the B1a RDMA microbenchmark + GLM-5.1 PD harness scripts.
Governance work item WI-2026-06-06-001 tracked alongside.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:04 +00:00
leaveletandClaude Opus 4.8 9a40e42384 fix(disagg/mooncake): notify prefill on decode-side abort
Port the decode->prefill half of upstream sgl-project/sglang #27372 (the worker
skip-Failed guard half landed in cb1a03f0a). In multi-prefill/multi-decode PD, a
decode-initiated abort frees the decode's KV pages back to the allocator, but the
prefill never learns of it (request_status is per-process), so the prefill keeps
RDMA-writing the remaining chunks into pages the decode may have reallocated to a
different live request -> KV corruption. The already-ported worker guard is a no-op
here because nothing sets the prefill room to Failed.

Now the decode receiver sends a 4-field b"ABORT" notification to every prefill peer
(on abort() and on the poll() WaitingForInput timeout, at most once); the prefill
marks the room Failed (so the worker guard fires) and replies b"ABORT_ACK". The
ABORT branch is handled before the unconditional waiting_req_bytes[3] decode in
bootstrap_thread (a 4-field ABORT would otherwise crash the else branch), and
ABORT_ACK before the 3-tuple unpack in the decode thread.

De-entangled from the staging/tracing infra this branch lacks. Component-tested
(test_mooncake_abort_protocol.py: format, send-once guard, abort wiring, ZMQ
round-trip). Cross-process corruption-window closure to be verified by the
multi-P/D PD harness.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:04 +00:00
leaveletandClaude Opus 4.8 c6c88c2617 fix(disagg): account KV transfer metrics by actually-sent bytes
Port upstream sgl-project/sglang #24416 (staging-free). Our transfer metrics
were computed from the full prompt length, so total_mb / speed_gb_s were
systematically over-reported on every prefix-cache hit and every CP shared-KV
per-rank page filter. Now each sender accumulates the actually-sent KV/state
indices and reports bytes = sent_pages * per-page item bytes via a new
KVTransferMetric returned by get_transfer_metric(); prefill consumes that
instead of estimating from len(origin_input_ids), and skips fake-bootstrap and
dummy-CP-rank senders (which transfer nothing).

Also route convert_to_duration() phase deltas through duration_between(), which
returns 0 when a phase timestamp is uninitialized, fixing nonsensical negative
durations in the time-stats log.

Unit-tested (test_kv_transfer_metrics.py, 10 cases) in the CUDA-13 container.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:04 +00:00
leaveletandClaude Opus 4.8 2524fe4c9d fix(disagg/mooncake): mark abort Failed and skip aborted chunks in worker
Port the abort-correctness pair from upstream sgl-project/sglang:
- #24522: abort() now calls kv_mgr.update_status(room, Failed) so the
  transfer worker and other pollers observe the abort, not only the local
  conclude_state.
- #27372 (guard half): the transfer worker skips a chunk whose room has
  already failed or been aborted, so it never transfers into KV pages that
  may have been reclaimed. This matters more once per-layer async transfer
  widens the abort-vs-in-flight window.

The decode->prefill ABORT/ABORT_ACK notification half of #27372 is deferred:
it interleaves with staging/tracing infra this branch lacks and needs the PD
test harness to verify, so porting it now would be unverified guesswork.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:04 +00:00
leaveletandClaude Opus 4.8 0cd5c38a9e fix(disagg/mooncake): clip CP/NSA state indices to the shorter side
Port upstream sgl-project/sglang #23323. Our state_type in [swa, nsa]
transfer path only handled len(prefill_state_indices) < len(dst), and even
then clipped prefill_state_indices (a no-op when prefill is shorter), and
never handled the prefill > dst case. Now clip prefill when it is longer and
clip the dst state indices when it is shorter, so the extra-pool transfer
never reads past prefill_state_indices or writes past dst_state_indices.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 09:51:04 +00:00
leaveletandClaude Opus 4.6 e5982dcceb fix: call _update_leaf_status in inc/dec_node_lock_ref to prevent phantom evictable nodes
inc_node_lock_ref/dec_node_lock_ref adjusted evictable_size_ but did
not call _update_leaf_status, so nodes could become "phantoms" —
counted in evictable_size_ but missing from evictable_leaves. Under
load with write-behind, this caused eviction to return 0 despite
large evictable_size_, leading to OOM:

  Available tokens: 778624 (evictable_size=706880)
  evict_result=(num_tokens_evicted=0)

The race: write-behind locks node X via inc_node_lock_ref (X stays
in evictable_leaves). A request path touches X via inc_lock_ref,
which calls _update_leaf_status and removes X from evictable_leaves.
Request finishes, dec_lock_ref keeps X out (lock_ref still >0).
Write ack calls dec_node_lock_ref dropping lock_ref to 0, but never
calls _update_leaf_status — X is permanently lost.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-05-27 01:52:41 +08:00
leaveletandClaude Opus 4.7 3d65944a22 Remove stale layer_transfer_counter prefetch guards (port from e293d4a39 fix #3)
CpSharedKVMlaPrefetcher.create and CpSharedKVIndexPrefetcher.create
returned None whenever ``token_to_kv_pool.layer_transfer_counter`` was
set, which permanently disables NSA indexer prefetch whenever HiCache
is active — even when no H2D transfer is in progress.

The guard is unnecessary: buffer getters synchronize via wait_until(),
and the prefetcher's stream calls wait_stream(current_stream) before
materialization. Removing it restores intended prefetch parallelism
under HiCache + NSA.

Only fix #3 of e293d4a39 is portable to this branch — fixes #1 (SHM
timeout) and #2 (SHM exception safety) target shm_allreduce.py and the
batched-CP coordination path in hiradix_cache.py, neither of which
exists here.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 00:40:39 +08:00
leaveletandClaude Opus 4.7 3186382be1 Drop env-gated CP HiCache page-index validator
The ported 44ba832 added _validate_cp_hicache_page_indices gated on
envs.SGLANG_DEBUG_HICACHE_VALIDATE, but that env was introduced by
97a9f850c (not on this branch). Calling _write_cp / load_cp crashed:

  AttributeError: 'Envs' object has no attribute 'SGLANG_DEBUG_HICACHE_VALIDATE'

The validator is purely defensive (page-alignment holds by construction
in HostKVCache.alloc and CpSharedKVLayout.logical_locs_to_physical), so
remove the function entirely along with its 4 call sites in _write_cp
and load_cp. The wrapping try/except blocks existed solely to free
allocations when validation raised; with no raising call inside, they
become dead and are removed too. Stale imports (envs,
validate_page_aligned_token_indices) dropped.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 00:15:35 +08:00
leaveletandClaude Opus 4.7 1d630def95 Fix CP HiCache load_cp owner-pattern mismatch (cache-hit corruption)
In CP=8 + NSA-shared-KV + HiCache disagg-prefill, cache-hit prefill produced
incoherent decode output. Cold prefill on CP was correct; pure CP without
HiCache was correct. The bug lived at the HiCache load_cp / device-alloc
interface.

Root cause: cache_controller.load_cp called the plain
mem_pool_device_allocator.alloc(logical_len), which returns logical pages
with no CP owner-pattern preservation. Cold prefill instead uses
alloc_extend_compute_owner with a zigzag owner pattern from
build_in_seq_page_compute_owners. The saved CpHiCacheNodeMetadata.owned_positions
records WHICH POSITIONS in the write-time alloc were owned by this rank. At
load time, those same positions are applied to a new alloc whose per-position
owner pattern is arbitrary -- each rank loads its host bytes into physical
slots whose corresponding logical page is owned by a DIFFERENT rank.
Attention's materialize_shared_token_kv_buffer reads from the owner's
physical slot, which was never loaded. Result: garbage.

Fix:
- CpHiCacheNodeMetadata gains two required fields: page_owners (int8 per
  logical page, identical on all CP ranks) and page_size. __post_init__
  validates; split() bisects page_owners by page index with a page-alignment
  check.
- _write_cp derives page_owners from device_indices (page-first slot of each
  page -> logical page id -> layout.owner_for_logical_pages) and stores in
  both metadata-construction sites (zero-owned and normal).
- New CPSharedPagedTokenToKVPoolAllocator.alloc_pages_with_owners() reuses
  _select_compute_owner_pages (with its tai-kernel fast path) and returns
  page-contiguous token locs whose per-page owner sequence equals the input.
- load_cp now concats page_owners across nodes_to_load and calls
  alloc_pages_with_owners. On None (lane exhausted) the caller hits the
  retry-with-eviction path; further failure returns None and degrades to
  cache miss. No silent fallback to plain alloc -- that recreated the bug.
- load_back retry path now calls _evict_for_compute_owner_lanes (module-top
  import) instead of plain evict(); this targets the deficit lane and gives
  the next alloc attempt a chance to satisfy it.
- envs import moved to module top in cache_controller.py per code-review
  feedback. Removed an over-defensive owned_check.all().item() in load_cp
  that would have re-introduced the host-sync anti-pattern 97a9f850c
  removed -- the invariant is already guaranteed by alloc_pages_with_owners.

Tests: 40 existing CpHiCacheNodeMetadata constructions migrated to pass the
new required fields. 9 new metadata tests (validators + split page-alignment).
10 new allocator tests in test_alloc_pages_with_owners.py covering input-order
preservation, lane exhaustion, release_pages fallback, debug-mode invariant.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 00:01:39 +08:00
leaveletandClaude Opus 4.7 d655fad040 Dispatch EAGLE radix bigram builder to tai-kernel
The pure-Python convert_to_bigram_key list comprehension runs on every
cache_finished_req / cache_unfinished_req of every EAGLE radix variant
(radix_cache, hiradix_cache, swa_radix_cache), with token-list lengths
that scale with prompt + output.  Scheduler profiles consistently flag it
as the largest mem_cache-side CPU hotspot.

This commit wires sglang.srt.mem_cache.utils.convert_to_bigram_key to
tai_kernel.radix.convert_to_bigram_key when the extension is importable,
falling back to the pure-Python implementation otherwise.  The tai-kernel
path uses a pybind11 module that calls the CPython C API directly
(PyTuple_New + PyTuple_SET_ITEM with ref-stealing) rather than rebuilding
the list comprehension's bytecode-level tuple allocations.  Measured 1.4x
at n=131k and up to 2.5x for n=1k on g0034 Python 3.11; allocator-bound
at large n because CPython's 2-tuple freelist already amortises the
construction.  The int64-packing follow-up that bypasses tuple allocation
entirely is parked as a separate work item.

Runtime safety:
- The dispatcher catches any first-call JIT compile / runtime failure,
  logs once, and falls through to the pure-Python path for the rest of
  the process — JIT failures must degrade rather than crash a serving
  loop.
- SGLANG_DISABLE_TAI_BIGRAM forces the Python path for bisecting.

Constraint: Output must be a real Python List[Tuple[int, int]] because
downstream radix dicts use the tuples as hashable keys.

Rejected: Pre-allocated tuple slab pool | CPython's per-interpreter 2-tuple
freelist already serves this case, and we cannot recycle tuples that
become radix-tree keys without changing the consumer.

Rejected: int64-packed keys this round | requires changes to RadixKey,
get_child_key_fn, key_match_fn, and EAGLE bigram detection; deserves its
own plan.

Confidence: high

Scope-risk: low

Directive: Keep _python_convert_to_bigram_key reachable; if the tai-kernel
path is ever removed, the EAGLE radix cache must continue to work
unchanged.

Tested: tai-kernel side validated on the cluster
(``python benchmark/radix/benchmark_convert_to_bigram_key.py --check``
prints byte-exact correctness then 1.4-2.5x speedup across sizes
128..131072 on g0034).  The new
``test/registered/unit/mem_cache/test_convert_to_bigram_key.py``
exercises both dispatch paths via SGLANG_DISABLE_TAI_BIGRAM patching.

Not-tested: End-to-end EAGLE serving accuracy + scheduler-time delta on
a real workload.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 23:41:17 +08:00
leavelet bacad1d498 fix: update leaf status for hicache 2026-05-22 14:09:48 +00:00
leaveletandClaude Opus 4.6 bd6e28f8ce fix: embed before CP split in nextn to prevent TP all_reduce shape mismatch
With CP=8 and dp_size=1, enable_dp_attention gets reset to False, so
VocabParallelEmbedding uses tensor_model_parallel_all_reduce (tp_size=8).
The CP local draft path was splitting tokens before embedding, giving each
rank a different local_tokens count. This caused an NCCL all_reduce shape
mismatch and a collective hang.

Move embed_tokens() before the CP split: embed on the full input (all ranks
see the same shape), then cp_split_and_rebuild_data the result. The decoder
layer still runs on CP-local tokens, preserving the CP performance benefit.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-05-22 11:21:09 +00:00
leaveletandClaude Opus 4.6 8ee249f953 fix: re-export log_cp_draft_shared_kv_debug from cp_shared_kv_runtime
nsa_indexer.py imports log_cp_draft_shared_kv_debug from
cp_shared_kv_runtime, but ad358d164 defined it in nsa/utils.py.
The missing symbol caused deepseek_v2.py import to fail silently
(registry.py catches and warns), falling back to transformers.py
which loads unsharded 256-expert MoE parameters → OOM on H200.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-05-22 07:42:25 +00:00
leavelet 663459d2ef feat(deepgemm): deepgemm jit and init in model_runner
1. fix and tune deepgemm
2026-04-08 06:05:06 +08:00
leavelet 2ae0237d3e fix(deepep): tuning and enhance deepep on single node
feat(deepep): enhance single node deepep
2026-04-08 05:36:31 +08:00
leavelet e72880b467 feat(eplb): add eplb warmup 2026-04-08 05:21:54 +08:00
leavelet 0bfa394aff [Fix]: HiCache hasher failed when EAGLE mode enabled (#12025) 2025-10-24 23:53:13 +08:00
leavelet 19ba16aa3d [Fix]: add missing device attribute to ChunkCache (#11493) 2025-10-12 20:49:59 -07:00