כתבה
arXiv cs.CL ·
שיפור עיבוד טפלות דרך אימון מיקודי
Reasoning Fine-Tuning Induces Persistent Latent Policy States
חוקרים גילו כי אימון מיקודי משפר את יכולתם של מודלים לשפה גדולים (LLM) לבצע טפלות. המחקר הראה כי האימון המיקודי גורם לשינויים פנימיים במודל, שמשפרים את יכולתו לבצע טפלות מורכבות.
תקציר מקורי באנגליתarXiv:2607.18532v1 Announce Type: new Abstract: Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain poorly understood. It is unclear whether reasoning fine-tuning improves local token-level competence or globally reorganizes how models structure inference over time. We address this question by modeling Chain-of-Thought reasoning as a switching dynamical system (SDS), in which internal representations evolve under discrete latent policy states. Our framework combines time-aware contrastive representation learning with discrete regime discovery to recover latent policies from activation trajectories. Across four benchmarks and model scales from 1.5B to 32B parameters, reasoning-fine
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית