כתבה
arXiv cs.CL ·
קריסת מגוון במשחקי LLM
When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play
חוקרים בדקו את השפעת הסיכול המודרך על מגוון התנהגות במשחקים. התוצאות הראו כי ייצור במצב תכנון תכופות מדכא מגוון פעולות. החוקרים הציעו שיטה חדשה להכשרה שמטרתה לשמר מגוון פעולות.
תקציר מקורי באנגליתarXiv:2607.19523v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks, but its effect on behavioral diversity in sequential decision-making remains under-explored. We study this question in a controlled suite of deterministic board games based on tic-tac-toe variants, where optimal actions are exactly computable and diversity can be measured directly. Across state-level evaluation, arena gameplay, and training trajectories, we find that reasoning-mode generation frequently suppresses action diversity without uniformly improving action accuracy. Furthermore, standard SFT improves accuracy but often induces premature diversity collapse, which exceeds what is minimally required by the accuracy-diversity tradeoff. We then
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית