In CP=8 + NSA-shared-KV + HiCache disagg-prefill, cache-hit prefill produced incoherent decode output. Cold prefill on CP was correct; pure CP without HiCache was correct. The bug lived at the HiCache load_cp / device-alloc interface. Root cause: cache_controller.load_cp called the plain mem_pool_device_allocator.alloc(logical_len), which returns logical pages with no CP owner-pattern preservation. Cold prefill instead uses alloc_extend_compute_owner with a zigzag owner pattern from build_in_seq_page_compute_owners. The saved CpHiCacheNodeMetadata.owned_positions records WHICH POSITIONS in the write-time alloc were owned by this rank. At load time, those same positions are applied to a new alloc whose per-position owner pattern is arbitrary -- each rank loads its host bytes into physical slots whose corresponding logical page is owned by a DIFFERENT rank. Attention's materialize_shared_token_kv_buffer reads from the owner's physical slot, which was never loaded. Result: garbage. Fix: - CpHiCacheNodeMetadata gains two required fields: page_owners (int8 per logical page, identical on all CP ranks) and page_size. __post_init__ validates; split() bisects page_owners by page index with a page-alignment check. - _write_cp derives page_owners from device_indices (page-first slot of each page -> logical page id -> layout.owner_for_logical_pages) and stores in both metadata-construction sites (zero-owned and normal). - New CPSharedPagedTokenToKVPoolAllocator.alloc_pages_with_owners() reuses _select_compute_owner_pages (with its tai-kernel fast path) and returns page-contiguous token locs whose per-page owner sequence equals the input. - load_cp now concats page_owners across nodes_to_load and calls alloc_pages_with_owners. On None (lane exhausted) the caller hits the retry-with-eviction path; further failure returns None and degrades to cache miss. No silent fallback to plain alloc -- that recreated the bug. - load_back retry path now calls _evict_for_compute_owner_lanes (module-top import) instead of plain evict(); this targets the deficit lane and gives the next alloc attempt a chance to satisfy it. - envs import moved to module top in cache_controller.py per code-review feedback. Removed an over-defensive owned_check.all().item() in load_cp that would have re-introduced the host-sync anti-pattern 97a9f850c removed -- the invariant is already guaranteed by alloc_pages_with_owners. Tests: 40 existing CpHiCacheNodeMetadata constructions migrated to pass the new required fields. 9 new metadata tests (validators + split page-alignment). 10 new allocator tests in test_alloc_pages_with_owners.py covering input-order preservation, lane exhaustion, release_pages fallback, debug-mode invariant. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Run Unit Tests
SGLang uses the built-in library unittest as the testing framework.
Test Backend Runtime
cd sglang/test/srt
# Run a single file
python3 test_srt_endpoint.py
# Run a single test
python3 test_srt_endpoint.py TestSRTEndpoint.test_simple_decode
# Run a suite with multiple files
python3 run_suite.py --suite per-commit
Test Frontend Language
cd sglang/test/lang
# Run a single file
python3 test_choices.py
Adding or Updating Tests in CI
- Create new test files under
test/srtortest/langdepending on the type of test. - For nightly tests, place them in
test/srt/nightly/. Use theNightlyBenchmarkRunnerhelper class innightly_utils.pyfor performance benchmarking tests. - Ensure they are referenced in the respective
run_suite.py(e.g.,test/srt/run_suite.py) so they are picked up in CI. For most small test cases, they can be added to theper-commit-1-gpusuite. Sort the test cases alphabetically by name. - Ensure you added
unittest.main()for unittest andsys.exit(pytest.main([__file__]))for pytest in the scripts. The CI run them viapython3 test_file.py. - The CI will run some suites such as
per-commit-1-gpu,per-commit-2-gpu, andnightly-1-gpuautomatically. If you need special setup or custom test groups, you may modify the workflows in.github/workflows/.
CI Registry System
Tests in test/registered/ use a registry-based CI system for flexible backend/schedule configuration.
Registration Functions
from sglang.test.ci.ci_register import (
register_cuda_ci,
register_amd_ci,
register_cpu_ci,
register_npu_ci,
)
# Per-commit test (small 1-gpu, runs on 5090)
register_cuda_ci(est_time=80, suite="stage-b-test-1-gpu-small")
# Per-commit test (large 1-gpu, runs on H100)
register_cuda_ci(est_time=120, suite="stage-b-test-1-gpu-large")
# Per-commit test (2-gpu)
register_cuda_ci(est_time=200, suite="stage-b-test-2-gpu-large")
# Nightly-only test
register_cuda_ci(est_time=200, suite="nightly-1-gpu", nightly=True)
# Multi-backend test
register_cuda_ci(est_time=80, suite="stage-b-test-1-gpu-small")
register_amd_ci(est_time=120, suite="stage-a-test-1-gpu-small-amd")
# Temporarily disabled test
register_cuda_ci(est_time=80, suite="stage-b-test-1-gpu-small", disabled="flaky - see #12345")
Choosing Between 1-GPU Suites (5090 vs H100)
When adding 1-GPU tests, choose the appropriate suite based on hardware compatibility:
| Suite | Runner | GPU | When to Use |
|---|---|---|---|
stage-a-test-1-gpu-small |
1-gpu-5090 |
RTX 5090 (32GB, SM120) | Stage A per-commit smoke on 5090 (CUDA) |
stage-a-test-1-gpu-small-amd |
AMD CI runners | ROCm | Stage A per-commit smoke (AMD) |
stage-b-test-1-gpu-small |
1-gpu-5090 |
RTX 5090 (32GB, SM120) | 5090-compatible tests (preferred) |
stage-b-test-1-gpu-large |
1-gpu-h100 |
H100 (80GB, SM90) | Large models or 5090-incompatible tests |
Use stage-b-test-1-gpu-small (5090) whenever possible - this is the preferred suite for most 1-GPU tests.
Use stage-b-test-1-gpu-large (H100) if ANY of these apply:
-
Architecture incompatibility (SM120/Blackwell):
- FA3 attention backend (requires SM≤90)
- MLA with FA3 backend
- FP8/MXFP4 quantization (not supported on SM120)
- Certain Triton kernels (shared memory limits)
-
Memory requirements:
- Models >30B params or large MoE
- Tests requiring >32GB VRAM
-
Known 5090 failures:
- Weight update/sync tests
- Certain spec decoding tests
If a test cannot run on 5090 due to any of the above, use stage-b-test-1-gpu-large which runs on H100.
Available Suites
Per-Commit (CUDA):
- Stage A:
stage-a-test-1-gpu-small(5090),stage-a-test-2,stage-a-test-cpu - Stage B:
stage-b-test-1-gpu-small(5090),stage-b-test-1-gpu-large(H100),stage-b-test-2-gpu-large - Stage C (4-GPU):
stage-c-test-4-gpu-h100,stage-c-test-4-gpu-b200,stage-c-test-4-gpu-gb200,stage-c-test-deepep-4-gpu-h100 - Stage C (8-GPU):
stage-c-test-8-gpu-h20,stage-c-test-8-gpu-h200,stage-c-test-8-gpu-b200,stage-c-test-deepep-8-gpu-h200
Per-Commit (AMD):
stage-a-test-1-gpu-small-amd,stage-b-test-1-gpu-small-amd,stage-b-test-2-gpu-large-amd
Nightly:
nightly-1-gpu,nightly-2-gpu,nightly-4-gpu,nightly-8-gpu, etc.
Running Tests with run_suite.py
# Run per-commit tests
python test/run_suite.py --hw cuda --suite stage-b-test-1-gpu-small
# Run nightly tests
python test/run_suite.py --hw cuda --suite nightly-1-gpu --nightly
# With auto-partitioning (for parallel CI jobs)
python test/run_suite.py --hw cuda --suite stage-b-test-1-gpu-small \
--auto-partition-id 0 --auto-partition-size 4
Writing Elegant Test Cases
- Learn from existing examples in sglang/test/srt.
- Reduce the test time by using smaller models and reusing the server for multiple test cases. Launching a server takes a lot of time.
- Use as few GPUs as possible. Do not run long tests with 8-gpu runners.
- If the test cases take too long, considering adding them to nightly tests instead of per-commit tests.
- Keep each test function focused on a single scenario or piece of functionality.
- Give tests descriptive names reflecting their purpose.
- Use robust assertions (e.g., assert, unittest methods) to validate outcomes.
- Clean up resources to avoid side effects and preserve test independence.
- Reduce the test time by using smaller models and reusing the server for multiple test cases.
Adding New Models to Nightly CI
- For text models: extend global model lists variables in
test_utils.py, or add more model lists - For vlms: extend the
MODEL_THRESHOLDSglobal dictionary intest/srt/nightly/test_vlms_mmmu_eval.py