כתבה
arXiv cs.CL ·
SuperValid: אימות OOD לקנה מידה כללי
SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling
SuperValid הוא כלי לאימות OOD שמשפר את היכולת לחזות ביצועים במטלות שונות. הוא עובד על ידי יצירת נתונים חדשים המבוססים על מיומנויות משותפות. ניסויים מראים ש-SuperValid משפר את הדיוק בחיזוי ביצועים.
תקציר מקורי באנגליתarXiv:2605.28179v2 Announce Type: replace Abstract: Scaling laws guide large language model training by relating compute to cross-entropy loss, and recent work further extends them to predict downstream benchmark performance. However, prior approaches face generalization limitations from two aspects: focusing on benchmark-level performance introduces scenario-specific artifacts, while relying on IID validation loss fails to track capability improvements when training distributions vary. In this work, we argue that downstream scaling should be studied at the capability level, which captures shared skill factors across related tasks while abstracting away benchmark-specific noise. We propose SuperValid, a framework that synthesizes OOD (out-of-distribution), capability-aligned validation dat
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית