כתבה
arXiv cs.AI ·
בדיקת פרוטו-אינטרוספקציה במודלים שפה מחוגרים
Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary
חוקרים בדיקת פרוטו-אינטרוספקציה במודלים שפה מחוגרים. המחקר מראה כי מודלים אלו יכולים לקרוא את איכות החישוב שלהם ולשפר את התוצאות. הניסויים בוצעו על מודל Ouro-RLTT והראו שיפורים משמעותיים באיכות התוצאות.
תקציר מקורי באנגליתarXiv:2607.18553v3 Announce Type: replace-cross Abstract: Can a language model read the quality of its ongoing computation, and can an external intervention turn that readout into better outcomes? We test both questions in a frozen 2.6B looped transformer, Ouro-RLTT. On GSM8K, a strict pre-answer probe excludes the answer region and gold value yet predicts success: hidden states plus length/log-probability features reach AUROC 0.797 versus 0.731 for those surface features alone (increment +0.066; task-clustered 95% CI [+0.021,+0.112]; 170 tasks). On Horizon Logic, a prospectively extended task-disjoint study gives an increment of +0.111 (CI [+0.056,+0.169]), independently replicated on the new cohort (+0.095) and robust to an adversarial malformed-sibling shortcut. Recurrence also moves ca
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית