כתבה
arXiv cs.CL ·
מדוע פיתוחי זמן-המבחן ולימוד מוקדם נכשלים בקביעת עמדה פרטית?
Where Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction?
פיתוחי זמן-המבחן ולימוד מוקדם לא עוזרים בקביעת עמדה פרטית. ניתן לשפר את התוצאות באמצעות שיטה חדשה.
תקציר מקורי באנגליתarXiv:2609.33155v2 Announce Type: replace Abstract: Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stance prediction remains unclear. We study this question by predicting a person's stance in a new discussion from their history. We evaluate widely used test-time scaling strategies and post-training methods, such as supervised fine-tuning and reinforcement learning, and identify four failure modes across generation, selection, and learning: (1) incorrect consensus, where repeated samples agree on the wrong stance; (2) selection failure, where generation covers the observed stance but selection misses it; (3) response overfitting, where supervised fine-tuning improves imitation but harms predictio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית