כתבה
arXiv cs.LG ·
מחקר מעמיק של דגמי שפה קטנים במשימות תפיסה רציונלית
A Systematic Study of Small Language Models on Abstract Reasoning Tasks
במחקר זה נבחנו דגמי שפה קטנים במשימות תפיסה רציונלית. נמצא כי דגמי שפה קטנים יכולים להשיג תוצאות טובות במשימות אלה, אך הם רגישים לשינויים בהקשר.
תקציר מקורי באנגליתarXiv:2610.08680v1 Announce Type: new Abstract: Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer. Across more than 1,000 runs, we profile decoder-only, encoder--decoder, and mixture-of-experts model families under supervised fine-tuning. We examine the efficiency and stability of skill acquisition, robustness beyond the training distribution, interactions with model family and task formulation, and layer-wise attention signatures that accompany behavio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית