כתבה
arXiv cs.LG ·
התאמה עצמית תקינה: מדוגמות שחור-קופסה לתגובות ערך-פולשניות
Adaptive Self-Consistency: From Black-Box Sampling to Distribution-Valued Feedback
המאמר עוסק בפיתוח טכניקה של התאמה עצמית, המשתמשת בדוגמות שחור-קופסה ותגובות ערך-פולשניות. הטכניקה מאפשרת זיהוי מודלי הערך העיקרי באופן יעיל. המאמר מציג גם אלגוריתם חדש, ASC-D, המשתמש בטכניקה זו ומציג תוצאות משופרות בהשוואה לטכניקות קיימות.
תקציר מקורי באנגליתarXiv:2609.38931v1 Announce Type: new Abstract: Self-consistency samples many reasoning trajectories and aggregates their final answers, treating the LLM as a black box that returns one answer per trajectory. Yet the final answer of each trajectory is sampled from a softmax vector that is available from the model's log-probabilities. We refer to this as the grey-box setting in which each trajectory reveals this answer distribution rather than a single draw from it. We formulate efficient inference in this setting as sequential mode identification with distribution-valued observations: sample trajectories one at a time and stop as soon as the LLM's modal answer is identified at a prescribed confidence level. We characterize the asymptotic stopping rate of mode identification with distributi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית