Files
sglang/test/registered/unit
laoyao0822andOmX 250fab291d Account for MQA logits in CP batch admission
CP shared-KV bs>1 admission already bounds request count, extend tokens,
cached tokens, and an estimated temporary buffer size. The estimate missed
the fp32 MQA logits temporary, whose peak grows with query rows times
context rows and can dominate high-cache-hit multi-request batches.

Add an MQA logits peak term to the CPU-only estimator and include it in
the layer-forward peak enforced by --cp-shared-kv-prefill-max-buffer-size.
When SGLANG_NSA_MQA_LOGITS_CHUNK_MAX_ROWS is set, admission estimates the
post-chunk peak using that row cap; otherwise it remains conservative and
assumes the full extend-row count.

Constraint: Scheduler admission must stay CPU-only and cannot query CUDA free memory.
Rejected: Add a separate scheduler limit for MQA logits | the existing max-buffer-size knob is the right aggregate admission budget.
Rejected: Use SGLANG_NSA_MQA_LOGITS_FREE_MEM_FRACTION in scheduler | that depends on runtime CUDA free memory and would make admission host-sync or stale.
Confidence: medium
Scope-risk: moderate
Directive: Keep the estimator conservative when chunk max rows is unset; do not rely on CUDA free-memory queries in scheduler admission.
Tested: Local py_compile for estimator, scheduler, schedule_policy, and estimator tests.
Tested: Local pytest test_cp_shared_kv_prefill_buffer_estimator.py: 5 passed.
Tested: Remote g0034 cjy-glm5-new py_compile and estimator pytest: 5 passed.
Not-tested: ETE scheduler admission under high-cache-hit bs>1 traffic.

Co-authored-by: OmX <omx@oh-my-codex.dev>
2026-06-11 03:25:22 +08:00
..

Unit Tests

Component-level tests that do not launch a server or load model weights. Tests can use CPU or GPU — the key criterion is no server process.

Quick Start

  1. Find the source file under python/sglang/srt/.
  2. Create the corresponding test here, mirroring the source tree:
    srt/mem_cache/radix_cache.py       →  unit/mem_cache/test_radix_cache.py
    srt/sampling/sampling_params.py    →  unit/sampling/test_sampling_params.py
    
  3. Register for CI at the top of the file (after imports, before test classes):
    from sglang.test.ci.ci_register import register_cpu_ci
    register_cpu_ci(est_time=5, suite="stage-a-test-cpu")
    # or: register_cuda_ci(est_time=10, suite="stage-b-test-1-gpu-small")
    
  4. Run locally:
    pytest test/registered/unit/ -v            # all unit tests
    pytest test/registered/unit/mem_cache/ -v  # one module
    
  5. Run with coverage:
    # summary
    pytest test/registered/unit/ --cov --cov-config=.coveragerc -v
    
    # PR incremental check (require ≥60% on changed lines)
    pytest test/registered/unit/ --cov --cov-config=.coveragerc --cov-report=xml
    diff-cover coverage.xml --compare-branch=origin/main --fail-under=60
    

Example

"""Unit tests for <module> — no server, no model loading."""

from sglang.test.ci.ci_register import register_cpu_ci

register_cpu_ci(est_time=5, suite="stage-a-test-cpu")

import unittest

from sglang.srt.<module> import TargetClass
from sglang.test.test_utils import CustomTestCase


class TestTargetClass(CustomTestCase):
    def test_basic_behavior(self):
        obj = TargetClass(...)
        self.assertEqual(obj.method(), expected)


if __name__ == "__main__":
    unittest.main()

Rules

  • No popen_launch_server() or Engine(...).
  • No model weight loading.
  • Use CustomTestCase (from sglang.test.test_utils, adds CI retry).
  • Use unittest.mock for dependencies that are expensive to construct.