Under CP, enable_storage=False, so StorageMetricsCollector (prefetch/backup tokens + bandwidth) never emits, and L2 pool usage was debug-log-only (_cp_maybe_log_lane_stats) while L3 was invisible. Add a CpHiCacheMetricsCollector covering both, gated on enable_metrics alone, per-rank via the tp_rank label. - CpL3Store.stats(): per-payload slot occupancy (used,total) + cumulative counters (spill_pages, reload_pages, gc_reclaimed, write_errors -- incremented on the bg threads) + instantaneous inflight (spill/reload) + the four bg queue depths. Counters are monotonic (NOT reset on clear()). - CpHiCacheMetricsCollector (metrics_collector.py): gauges for instantaneous state (sglang:cp_l2_used/total_pages, cp_l3_used/total_slots, cp_l3_inflight_*, cp_l3_queue_depth) + counters for cumulative work (cp_l2_objects_committed/evicted_total, cp_l3_spill/reload_pages_total, cp_l3_gc_reclaimed_total, cp_l3_write_errors_total). Counters fed via a monotonic .inc(delta) off the store's cumulative totals. - hiradix: create the collector in _maybe_init_cp_l3 (once, per-rank labels from cache_controller); push periodically (~1s, per-rank wall-clock -- metrics need no rank-uniformity) from check_hicache_events via _cp_maybe_collect_metrics, reading allocator.occupancy_by_payload()/stats() + store.stats(). Off the data path. store.stats() unit-tested (occupancy + spill/reload counters; 10 store tests green). The collector + wiring are correct-by-pattern (mirror StorageMetricsCollector) and validate on the live server (3.4; prometheus_client not in the unit env). Directly useful for 3.4: watch L2/L3 fill + GC + reload hits under cachebench. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Registered Tests
Tests under this directory are auto-discovered by run_suite.py via CI registration decorators.
Where Should I Put My New Test?
No server / engine launch required
| What you're testing | Directory | Requires |
|---|---|---|
| Component logic in isolation (cache, scheduler, config, parser, etc.) | unit/<module>/ |
CPU or GPU |
| CUDA kernel correctness | kernels/ |
GPU |
Server / engine launch required (E2E)
| What you're testing | Directory | Requires |
|---|---|---|
| Model inference correctness | models/, 4-gpu-models/, 8-gpu-models/ |
GPU |
| Feature-specific (OpenAI API, LoRA, speculative, distributed, VLM, etc.) | openai_server/, lora/, spec/, distributed/, ... |
GPU |
| Benchmarks (performance, accuracy, stress) | benchmark/ |
GPU |
| Platform-specific | amd/, ascend/ |
Vendor GPU |
See unit/README.md for unit test conventions.