The publish-variant staging exchange measured ~equal to the default compact-current AR (90.4 vs 86.5 ms/batch on the traced scenario): the publish copy and the barrier serialized behind the 0.65 ms prefix gather ate the transport win the isolated current exchange showed (0.196 vs 0.354 ms). Fix the structure instead of the copy: current rows are now written straight INTO the staging — the fill kernels take their write destinations solely from page_inverse, so a per-batch staging-remapped page inverse on the plan retargets them with zero kernel changes — then cp_symm_barrier, then ONE slot-dense gather covers prefix pages (pool pointers) and ALL current pages (staging pointers, including this rank's own) through a concatenated 2*cp pointer table where current slots carry owner = cp_size + writer and src = staging slot. No publish copy, no prefix pre-gather, no second gather. The fused fill's loc outputs are dense-geometry-bound, so the token-KV path computes mixed_locs/staging row indices once per batch (they are layer-invariant) and the per-layer fill collapses to a single index_copy_ into the zeroed staging span. Benchmark (g0033 8xH200, byte-exact, idle-checked): 62.8 ms/batch vs 84.4 default Step A (-26%) and 60.8 ideal; publish variant was 88.0. 151 unit tests; 8-rank GPU byte-exactness vs v2 across 8 layers, arena on and off. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Unit Tests
Component-level tests that do not launch a server or load model weights. Tests can use CPU or GPU — the key criterion is no server process.
Quick Start
- Find the source file under
python/sglang/srt/. - Create the corresponding test here, mirroring the source tree:
srt/mem_cache/radix_cache.py → unit/mem_cache/test_radix_cache.py srt/sampling/sampling_params.py → unit/sampling/test_sampling_params.py - Register for CI at the top of the file (after imports, before test classes):
from sglang.test.ci.ci_register import register_cpu_ci register_cpu_ci(est_time=5, suite="stage-a-test-cpu") # or: register_cuda_ci(est_time=10, suite="stage-b-test-1-gpu-small") - Run locally:
pytest test/registered/unit/ -v # all unit tests pytest test/registered/unit/mem_cache/ -v # one module - Run with coverage:
# summary pytest test/registered/unit/ --cov --cov-config=.coveragerc -v # PR incremental check (require ≥60% on changed lines) pytest test/registered/unit/ --cov --cov-config=.coveragerc --cov-report=xml diff-cover coverage.xml --compare-branch=origin/main --fail-under=60
Example
"""Unit tests for <module> — no server, no model loading."""
from sglang.test.ci.ci_register import register_cpu_ci
register_cpu_ci(est_time=5, suite="stage-a-test-cpu")
import unittest
from sglang.srt.<module> import TargetClass
from sglang.test.test_utils import CustomTestCase
class TestTargetClass(CustomTestCase):
def test_basic_behavior(self):
obj = TargetClass(...)
self.assertEqual(obj.method(), expected)
if __name__ == "__main__":
unittest.main()
Rules
- No
popen_launch_server()orEngine(...). - No model weight loading.
- Use
CustomTestCase(fromsglang.test.test_utils, adds CI retry). - Use
unittest.mockfor dependencies that are expensive to construct.