כתבה
arXiv cs.AI ·
לא רק ציונים: תקן המאירות לאודיטוריית AI
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
תקן המאירות (DCP) מאפשר לאודיטוריית AI לבדוק טענות של סוכני AI על ידי בדיקות מבוצעות. התקן כולל שלושה חלקים: Gate 1, Gate 2 ו-Gate 3. Gate 1 ו-Gate 2 נועדו לבדוק את יעילות הסוכן, ו-Gate 3 נועד לבדוק את השפעת הפידבק הנאמן. DCP נועד לספק תקן רשמי לאודיטוריית AI.
תקציר מקורי באנגליתarXiv:2609.09219v1 Announce Type: cross Abstract: AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement on sealed evaluation. Gate 2 gives matched agents the registered starting information and observed Web content while withholding the target research history. Every valid method reaching the numerical target supplies a recovery witness and triggers the Core veto. DCP Core requires adequate controls, zero observed recoveries, and a finite-sample bound on recovery in one fresh registered episode. Optional Gate 3 measures the average effect of truthful feedback relative to a spec
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית