Route CP shared MLA store through TAI fused kernels without runtime spam
The shared-KV prefill path now optionally calls tai_kernel.nsa_prefill.fused_store_mla_kv before falling back to logical_locs_to_physical plus set_mla_kv_buffer. The fast path supports packed FP8 and BF16/FP16 direct KV buffers, while debug mode and kernel failures still preserve the existing fallback behavior. Success logging was removed after path verification because per-layer/per-rank logs are too noisy in normal server runs. Constraint: Runtime must remain safe when tai-kernel is absent or debug checks are enabled Rejected: Keep success logs permanently | floods prefill logs once every rank/layer starts using the fast path Confidence: high Scope-risk: moderate Directive: Keep fallback warnings; do not re-add per-layer success logs outside explicit debug instrumentation Tested: g0034 container python -m py_compile python/sglang/srt/layers/attention/nsa/cp_shared_kv_runtime.py Tested: g0034 container PYTHONPATH=python pytest -q test/registered/unit/mem_cache/test_cp_shared_kv_runtime.py -q (40 passed) Not-tested: Full multi-node PD server throughput after log removal
This commit is contained in:
@@ -207,6 +207,7 @@ class Envs:
|
||||
SGLANG_DEBUG_MOE_SORT_NVTX = EnvBool(False)
|
||||
SGLANG_CP_SHARED_KV_CURRENT_REUSE = EnvBool(False)
|
||||
SGLANG_CP_SHARED_KV_USE_TAI_MATERIALIZE = EnvBool(False)
|
||||
SGLANG_CP_SHARED_KV_FUSED_MLA_STORE = EnvBool(False)
|
||||
SGLANG_CP_SHARED_KV_ENABLE_MLA_PREFETCH = EnvBool(False)
|
||||
SGLANG_CP_SHARED_KV_LOG_MLA_PREFETCH = EnvBool(False)
|
||||
SGLANG_TEST_REQUEST_TIME_STATS = EnvBool(False)
|
||||
|
||||
Reference in New Issue
Block a user