כתבה
arXiv cs.AI ·
ג'נרליזציה היא יציבות, ולא דיוק: מבחן רב-ממדי של LLMs
Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
במאמר זה, המחברים מציגים דרך חדשה לבדיקת ג'נרליזציה של LLMs, המתמקדת ביציבות ולא בדיוק. הם מציגים פרקטיקה חדשה, SAGO, שמדדת את השינויים בהתנהגות המודל כשהוא נתקל באותו הקלט בגרסאות שונות. המחברים מראים שאף גם המודלים המוכרים ביותר נוטים לאינסטביליות ג'נרליזציה: אין מודל שג'נרליזציה באופן יושר, והממדים השונים של התנהגות המודל חושפים נקודות עתירות.
תקציר מקורי באנגליתarXiv:2610.01428v1 Announce Type: cross Abstract: Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, which conflates robustness with overall benchmark performance. In this work, we show generalization evaluation at the level of individual examples, across multiple input variants, and across different aspects of model behavior, focusing on variability rather than reducing performance to a score that can be improved through narrow training or other ways that obfuscate generalization evaluation. Following this view, we introduce the Stability-Aware
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית