כתבה
arXiv cs.CL ·
Where a Model Sends Its Own Repeated Token
תקציר מקורי באנגליתarXiv:2609.31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We take that input and ask where the model goes when it does not. For each token t, read argmax p(. | t, t) in one forward pass; the result is a map on the whole vocabulary, with two halves. The first -- which tokens are fixed points -- is partially anticipated, and we report it as a failed estimand: the natural distance on it is 83% cardinality, separates a corpus manipulation by two bits in 3471 against a precision floor of zero, and attributes families at 0.5833. The second half, where the map sends tokens that are not
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית