The per-layer KV transfer registration hardcoded chunk_page_start=0 when filtering CP-shared-KV owned pages. The CP filter's second return (`positions`) are absolute full-sequence page positions built from chunk_page_start, and the transfer indexes the FULL-request dst_kv_indices by those absolute positions (mirroring the monolithic send(), which passes chunk_page_start=index_slice.start — the cumulative page offset). With start=0, chunk N>0's positions were chunk-local, so its KV was written onto chunk 0's decode pages, corrupting the decode output. Non-chunked requests (single chunk, start=0) were unaffected, matching the observed symptom (non-chunked byte-identical, chunked garbage). Fix: chunk_page_start = chunk_key // page_size, where chunk_key is the chunk's start_send_idx (page-aligned), making it exactly the monolithic index_slice.start. Verified: opus first-principles code audit; empirical mapping-invariant on the deployed modules (per-chunk == whole-request for all 8 CP ranks; old start=0 sends chunk1 to chunk0's dst); 2 new regression tests (TestChunkedDstMapping). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Unit Tests
Component-level tests that do not launch a server or load model weights. Tests can use CPU or GPU — the key criterion is no server process.
Quick Start
- Find the source file under
python/sglang/srt/. - Create the corresponding test here, mirroring the source tree:
srt/mem_cache/radix_cache.py → unit/mem_cache/test_radix_cache.py srt/sampling/sampling_params.py → unit/sampling/test_sampling_params.py - Register for CI at the top of the file (after imports, before test classes):
from sglang.test.ci.ci_register import register_cpu_ci register_cpu_ci(est_time=5, suite="stage-a-test-cpu") # or: register_cuda_ci(est_time=10, suite="stage-b-test-1-gpu-small") - Run locally:
pytest test/registered/unit/ -v # all unit tests pytest test/registered/unit/mem_cache/ -v # one module - Run with coverage:
# summary pytest test/registered/unit/ --cov --cov-config=.coveragerc -v # PR incremental check (require ≥60% on changed lines) pytest test/registered/unit/ --cov --cov-config=.coveragerc --cov-report=xml diff-cover coverage.xml --compare-branch=origin/main --fail-under=60
Example
"""Unit tests for <module> — no server, no model loading."""
from sglang.test.ci.ci_register import register_cpu_ci
register_cpu_ci(est_time=5, suite="stage-a-test-cpu")
import unittest
from sglang.srt.<module> import TargetClass
from sglang.test.test_utils import CustomTestCase
class TestTargetClass(CustomTestCase):
def test_basic_behavior(self):
obj = TargetClass(...)
self.assertEqual(obj.method(), expected)
if __name__ == "__main__":
unittest.main()
Rules
- No
popen_launch_server()orEngine(...). - No model weight loading.
- Use
CustomTestCase(fromsglang.test.test_utils, adds CI retry). - Use
unittest.mockfor dependencies that are expensive to construct.