The host-pool size check used psutil.virtual_memory().available -- regular RAM -- unconditionally. But
hugepages are carved OUT of regular RAM, so reserving e.g. 1 TB of 2M hugepages drops psutil.available
by ~1 TB, and the check then spuriously fails ("Not enough host memory ... only have 897 GB free") even
though the slab is allocated from the hugepages, not regular RAM. It was inspecting the wrong pool.
Two fixes:
- memory_pool_host.py: gate the psutil check on the DEFAULT (regular-malloc) allocator. A custom
host_tensor_allocator (the CP shared-L2 slab over hugetlbfs) owns its own capacity check, so skip the
regular-RAM check for it.
- cp_shared_l2_pool.py: add cp_shared_l2_hugetlbfs_free_bytes() + check_cp_shared_l2_hugetlbfs_capacity()
(statvfs the hugetlbfs mount: f_frsize = hugepage size, f_bavail = free hugepages) and call it in
create_cp_shared_host_slab before ftruncate/mmap. Now insufficient hugepages fail LOUD with a clean,
actionable message ("insufficient hugepages in /dev/hugepages: need X GB (N x 2M pages) but only Y GB
free; reserve more or reduce --hicache-size") instead of a cryptic mmap ENOMEM. Per-slab + sequential,
so each payload's slab sees the free remaining after the prior ones.
Unit-tested (capacity check: tiny fits, over-capacity fails loud). Pre-existing 2b L2-pooling code (not L3).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Unit Tests
Component-level tests that do not launch a server or load model weights. Tests can use CPU or GPU — the key criterion is no server process.
Quick Start
- Find the source file under
python/sglang/srt/. - Create the corresponding test here, mirroring the source tree:
srt/mem_cache/radix_cache.py → unit/mem_cache/test_radix_cache.py srt/sampling/sampling_params.py → unit/sampling/test_sampling_params.py - Register for CI at the top of the file (after imports, before test classes):
from sglang.test.ci.ci_register import register_cpu_ci register_cpu_ci(est_time=5, suite="stage-a-test-cpu") # or: register_cuda_ci(est_time=10, suite="stage-b-test-1-gpu-small") - Run locally:
pytest test/registered/unit/ -v # all unit tests pytest test/registered/unit/mem_cache/ -v # one module - Run with coverage:
# summary pytest test/registered/unit/ --cov --cov-config=.coveragerc -v # PR incremental check (require ≥60% on changed lines) pytest test/registered/unit/ --cov --cov-config=.coveragerc --cov-report=xml diff-cover coverage.xml --compare-branch=origin/main --fail-under=60
Example
"""Unit tests for <module> — no server, no model loading."""
from sglang.test.ci.ci_register import register_cpu_ci
register_cpu_ci(est_time=5, suite="stage-a-test-cpu")
import unittest
from sglang.srt.<module> import TargetClass
from sglang.test.test_utils import CustomTestCase
class TestTargetClass(CustomTestCase):
def test_basic_behavior(self):
obj = TargetClass(...)
self.assertEqual(obj.method(), expected)
if __name__ == "__main__":
unittest.main()
Rules
- No
popen_launch_server()orEngine(...). - No model weight loading.
- Use
CustomTestCase(fromsglang.test.test_utils, adds CI retry). - Use
unittest.mockfor dependencies that are expensive to construct.