יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Kepler: דגימות עולם ניתנות לבדיקה ל-ARC-AGI-3

Kepler: Auditable World Models for ARC-AGI-3
Kepler היא דגימה עולם ניתנת לבדיקה שמייצגת השערות כדגימות עולם ניתנות לבדיקה. היא נבחנה ב-ARC-AGI-3, שבה נבדקים agents באזורי תקיפה שבהם צריך להיות נתונים כדי ללמוד. Kepler השיגה 100.00 RHAE ב-25 משחקים ציבוריים, וב-181 מ-183 מהשלבים הסופיים, הניסיון האחרון של Opus השתמש ב-8,256 פעולות של האזור, כולל 7,292 בשלבים שנכללו בסקור.
תקציר מקורי באנגליתarXiv:2610.00834v1 Announce Type: new Abstract: ARC-AGI-3 evaluates agents in interactive environments whose rules and objectives must be inferred from observation. We present Kepler, an open-source harness that represents hypotheses as executable world models and validates them through retrospective transition checks and conditional prediction checks. Under one frozen Claude Opus 5 configuration, Kepler obtained a server-verified 100.00 RHAE on all 25 public games, with no per-game model selection or score-conditioned reruns. On 181 of 183 completed levels, the final Opus attempt used no more actions than the corresponding median-human baseline. The retained board runs used 8,256 environment actions, of which 7,292 occurred in scored levels. Retained local provider-session records yield 8
קרא במקור המקורי