יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

עבר מעבר לביטחון

Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning
TASCO הוא כלי חדש לשיפור יכולות ההיגיון של מודלי שפה גדולים. הוא משתמש באותות מודל מונחים כדי להדריך את המודלים למצבי היגיון בטוחים יותר. TASCO משפר את דיוק ההיגיון ואת יעילות הטוקנים במודלים שונים.
תקציר מקורי באנגליתarXiv:2609.11393v1 Announce Type: new Abstract: Test-time adaptation has emerged as a lightweight alternative to costly post-training for improving the reasoning capabilities of Large Language Models (LLMs) on downstream tasks. Predictive entropy provides a model-derived signal for such adaptation, guiding models toward higher-confidence reasoning states without external verifiers or reward models. However, higher confidence does not necessarily imply correctness, as LLMs may remain highly confident along incorrect reasoning trajectories. We observe that high-confidence reasoning is more likely to be correct when confidence remains stable under local perturbations. Based on this observation, we propose Test-Time Adaptation via Stability-Aware Confidence Optimization (TASCO), a framework th
קרא במקור המקורי