יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

כאשר הניסויים לא מתארגנים: פרויקטיביליות בבדיקת AI

When benchmark inferences do not compose: Projectibility in AI evaluation
תוצאות ניסויי AI לא נוטות להוביל לדברים חשובים. המבדקים מכלילים ומגזימים, אך התורות המתמקדות בעדות דורשות ראיות לכל טענה. המאמר זה עוסק בבעיה האפיסטמולוגית של קשרים חזקים שלא מתקשרים. המאמר זה טוען כי תוצאות ניסויי AI לא נוטות להוביל לדברים חשובים. המבדקים מכלילים ומגזימים, אך התורות המתמקדות בעדות דורשות ראיות לכל טענה.
תקציר מקורי באנגליתarXiv:2607.26159v1 Announce Type: new Abstract: An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from obs
קרא במקור המקורי