יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

תשובות נכונות, תווי-עקב לא תקינים: מה ידעי-בית יכול ללמד אותנו על תווי-עקב של רשתות

Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces
תווי-עקב של רשתות, המציגים את דרך החשיבה שלהן, לא תמיד תקינים, גם אם התשובה היא נכונה. כך נמצא במחקר חדש, המשתמש בידעי-בית כדי לבדוק את תקינות התווי-עקב. המחקר גילה ש-31.6% מהתשובות הנכונות היו עם תווי-עקב לא תקינים. כמו כן, נמצא ששינויים קלים בתווי-עקב, כגון החלפת תווים, לא השפיעו על תקינות התשובה. המחקר טוען שאלו התגליות יכולות להיות חשובות לבטיחות הרשתות.
תקציר מקורי באנגליתarXiv:2609.38107v1 Announce Type: new Abstract: Chain-of-thought traces are widely read as records of how models reach their answers, informing debugging, agent auditing, and claims about reasoning. Testing this interpretation is difficult because natural-language thinking traces are rarely mechanically verifiable. We revisit it in iGSM, a synthetic grade-school mathematics benchmark designed to study thinking traces and used to support claims of learned reasoning and planning. Crucially, iGSM exposes the exact quantities and dependencies that a correct solution should use, allowing generated traces to be checked programmatically step by step and enabling us to test whether correct answers are reliably accompanied by valid traces. We first evaluate models trained exclusively on valid, mini
קרא במקור המקורי