EAGLE verification can pass torch scalar tensors through the speculative tree traversal before calling grammar backends. xgrammar/tvm_ffi requires Python int token ids, so normalize traversal indices and draft token ids at the boundary while leaving tensor storage unchanged. Constraint: xgrammar/tvm_ffi rejects torch scalar tensors for GrammarMatcher.accept_token Rejected: Coerce tokens inside every grammar backend | the invalid value originates in speculative traversal and should be fixed before backend dispatch Confidence: high Scope-risk: narrow Directive: Keep grammar backend calls scalar-Python typed; do not pass torch scalar tensors through accept_token Tested: remote g0034 cjy-glm5-new PYTHONPATH=python python -m pytest -q test/registered/unit/speculative/test_spec_utils.py test/registered/unit/configs/test_nsa_index_layers.py test/registered/unit/models/test_deepseek_index_skip_weight_loading.py -> 19 passed Tested: remote g0034 cjy-glm5-new py_compile for modified runtime files Not-tested: live decode replay after this commit
1.4 KiB
1.4 KiB