Fail fast when CP compose would hide dense fallback

CP shared-KV bs>1 compose must not silently fall back to dense full-buffer collectives when CUDA TAI materialize is expected. The fallback masks both correctness contract drift and severe synchronization/communication regressions, especially while comparing the symm path with the older IPC path.\n\nThis keeps CPU/unit-test fallback available, but makes production CUDA+TAI runs raise an explicit compose_v2 fail-fast for token-KV and index dense fallback. It also records the symm-vs-IPC comparison contract so barrier and collective counts are evaluated alongside elapsed time.\n\nConstraint: Production cache-hit-heavy bs>1 paths must expose unexpected dense collectives instead of silently taking them.\nRejected: Cherry-pick the old IPC branch wholesale | it conflicts with the symm compose design and would mix two transport protocols before benchmarking.\nRejected: Allow dense fallback with warning only | warning can be missed and still corrupts performance conclusions.\nConfidence: medium\nScope-risk: moderate\nDirective: Do not re-enable dense full fallback in CUDA+TAI compose paths without a benchmark proving it is intentional and a correctness test covering cache-hit bs>1.\nTested: python -m py_compile for cp_shared_kv_runtime.py and test_cp_shared_kv_runtime.py; git diff --check.\nNot-tested: Remote container pytest/ETE; local pytest is not reliable in this workspace because dependencies such as orjson are missing.
This commit is contained in:
laoyao0822
2026-06-12 23:57:40 +08:00
parent 9d65bdba95
commit 2387787ebc
3 changed files with 174 additions and 0 deletions
@@ -734,3 +734,53 @@ P7 GSM8K/replay/Nsight 验证
```
不要先做 P6 再修 P2/P3。0SM kernel 只解决 transport 代价,不解决 bs>1 prefix/current slot 语义;如果语义仍是 scalar prefix,kernel 越快只会越快地产生错误。
## 2026-06-12 追加:symm compose 与旧 IPC kernel 的对比口径
`symm-syh` 分支的 compose/symm kernel 可能比旧 CUDA IPC current-staging kernel 单次更快,但不能只比较 kernel elapsed time。需要把同步点作为一等指标,否则可能出现“单 kernel 快,但每层/每 buffer 多一次同步,ETE 更慢”的误判。
对比 benchmark / Nsight trace 必须至少拆分以下事件:
1. **IPC capability agreement**
- `_agreed_tai_ipc_peer_ptrs()` 里的 group agreement / all-reduce。
- 应确认是每个 pool tensor 一次,还是每个 forward/layer 重复触发。
2. **symm barrier**
- `cp_symm_barrier()` 调用次数、耗时、等待方差。
- 需要按 token KV / index buffer 分开统计。
- 当前 tai-kernel 实现是 1 个 block / 1 个 warp 的 CUDA spin barrier,
不是 copy-engine/0SM 路径;它占用很少 SM,但会在当前 stream 上形成
明确同步点。判断 symm 是否优于旧 IPC 时,必须把这个 barrier 的次数和
rank 间等待方差算进去。
3. **symm mega gather**
- `materialize_cuda_ipc_peer_pages_slot_dense()` 在 symm combined ptr table 上的耗时。
- 统计 prefix pages、current pages、request 数、dense pages。
4. **compact current reduce**
- symm 未开启或不可用时的 `_reduce_current_pages_compact()`。
- 这是 compact current collective,不是 dense full fallback;但仍会同步/占用通信资源,需要单独计数。
5. **dense full fallback reduce**
- `_all_reduce_materialized_buffer(... v2_full ...)`。
- 生产 CUDA + TAI materialize 开启时不允许静默发生;应 fail-fast 暴露 `CP_SHARED_KV_FAIL_FAST][compose_v2]`。
6. **CPU descriptor / plan 成本**
- `get_or_build_compose_plan()` cache hit/miss。
- per-forward 是否复用 descriptor;不要把一次性 build 成本误算到每层 steady state。
建议 benchmark 矩阵:
- dtype:bf16 / fp8_e4m3;
- batch size:1, 2, 5, 10;
- extend:1k, 2k, 10k, 40k, 65k;
- cached/prefix:100k, 200k, 300k;
- case:cache-hit partial-current、current-only、multi-request shared prefix;
- mode:legacy dense fallback(只作为基线,不允许生产静默)、old CUDA IPC current staging、symm staging、symm prefetch。
结论标准:保留默认路径必须同时满足:
- 无 dense full fallback;
- 同步点数量不高于旧 IPC 路径,或同步耗时被更少 kernel/更高带宽抵消;
- ETE replay 在 cache-hit-heavy 短 extend 场景提升,而不是只在 micro benchmark 提升;
- GSM8K cache-hit 二轮精度不掉点。