יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אימות טענות, לא ציונים

Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents
חוקרים מציגים שיטה לאימות טענות של סוכנים מודולריים. השיטה בודקת ראיות ולא ציונים. היא משמשת להערכת השיפורים בסוכנים וזיהוי המרכיבים שאיבדו ערך.
תקציר מקורי באנגליתarXiv:2610.01348v1 Announce Type: new Abstract: When developers change one component of an agent, such as its controller, a learned model or its verifier, they usually judge the change by an aggregate task score. That score cannot tell whether improvement was attainable, which component lost value, or what the agent's own checks certify. We introduce a claim-specific verification audit for modular agents that plan, act, check and refine. Instead of scoring the agent, the audit scores the evidence: each conclusion is recorded with the evidence behind it, one of four verdicts (supported, unsupported, unresolved or not evaluated) and the boundary within which it holds. Three tools supply that evidence. Oracle policies measure attainable improvement under an explicitly stated action set, so th
קרא במקור המקורי