Files
sglang/test/registered
laoyao0822 2de822661c Avoid decode-step syncs in min-new-token penalties
The old min_new_tokens penalizer updated logits through boolean-mask indexing. That indexing is data-dependent and can force synchronization on the decode hot path.

Use an elementwise torch.where followed by inplace add so the operation stays tensorized and avoids the mask-index update path.

Constraint: Keep the numeric behavior for active rows equivalent without multiplying by zero, which would turn -inf penalties into NaN.

Rejected: Expanding the mask and assigning through logits[mask] | this is the synchronization pattern being removed.

Confidence: high

Scope-risk: narrow

Directive: Do not reintroduce boolean-mask writes in decode-step penalty paths without profiling the synchronization behavior.

Tested: RED/GREEN local pytest test/registered/unit/sampling/test_min_new_tokens_penalizer.py

Tested: RED/GREEN remote pytest in cjy-glm5-new for min_new_tokens penalizer together with related regression tests

Tested: git diff --check; py_compile for min_new_tokens.py

Not-tested: Full decode throughput benchmark
2026-06-29 03:16:15 +08:00
..
2026-03-21 17:10:35 +08:00
2025-11-26 13:24:39 -08:00

Registered Tests

Tests under this directory are auto-discovered by run_suite.py via CI registration decorators.

Where Should I Put My New Test?

No server / engine launch required

What you're testing Directory Requires
Component logic in isolation (cache, scheduler, config, parser, etc.) unit/<module>/ CPU or GPU
CUDA kernel correctness kernels/ GPU

Server / engine launch required (E2E)

What you're testing Directory Requires
Model inference correctness models/, 4-gpu-models/, 8-gpu-models/ GPU
Feature-specific (OpenAI API, LoRA, speculative, distributed, VLM, etc.) openai_server/, lora/, spec/, distributed/, ... GPU
Benchmarks (performance, accuracy, stress) benchmark/ GPU
Platform-specific amd/, ascend/ Vendor GPU

See unit/README.md for unit test conventions.