E2e caught one request wedged FOREVER in disagg_prefill_inflight_queue (probes 21s apart with zero traffic both showed #inflight-req: 1, no reap/timeout warnings ever logged). Mechanism, established by reading the full state machine: prefill Success is set locally by the transfer worker on the LAST chunk; if the decode peer is torn down between the handshake and the prefill's final send(), add_transfer_request silently drops the chunk (no transfer destinations) — Success becomes unreachable. The only external rescue, the decode ABORT notification, is best-effort (silently swallowed on send error, no-op if it races the room registration), there is no prefill-side heartbeat of decode sessions, and the sender's only timeout covers Bootstrapping — the inflight queue itself has no liveness bound. The orphan pins the request's KV pages and rides every poll collective. Two fixes, both reaped through the existing Failed branch via the CP/TP MIN-reduce poll consensus (Failed=0 wins, so one rank concluding flips every rank together — rank-uniform by construction): - add_transfer_request: a room with no transfer destinations that is NOT already Success (the dummy-rank handshake marking) now concludes Failed loudly instead of dropping the chunk silently. - Inflight residency timeout: entries stuck in a non-terminal poll state past SGLANG_DISAGGREGATION_INFLIGHT_TIMEOUT (default 300s, matching the sibling BOOTSTRAP/WAITING timeouts) get sender.abort() and reap on the next poll. Covers what the hardening cannot: lost ABORT datagrams, decode crashes. Known sibling gaps left for follow-up: the decode transfer queue has no Transferring liveness bound, and an abort that matches no queue is still a silent no-op (much narrower race than first thought — work requests are ordered before control requests within a tick). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Registered Tests
Tests under this directory are auto-discovered by run_suite.py via CI registration decorators.
Where Should I Put My New Test?
No server / engine launch required
| What you're testing | Directory | Requires |
|---|---|---|
| Component logic in isolation (cache, scheduler, config, parser, etc.) | unit/<module>/ |
CPU or GPU |
| CUDA kernel correctness | kernels/ |
GPU |
Server / engine launch required (E2E)
| What you're testing | Directory | Requires |
|---|---|---|
| Model inference correctness | models/, 4-gpu-models/, 8-gpu-models/ |
GPU |
| Feature-specific (OpenAI API, LoRA, speculative, distributed, VLM, etc.) | openai_server/, lora/, spec/, distributed/, ... |
GPU |
| Benchmarks (performance, accuracy, stress) | benchmark/ |
GPU |
| Platform-specific | amd/, ascend/ |
Vendor GPU |
See unit/README.md for unit test conventions.