יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

CARGO: חיפוש-מודלי וביקורת-מבוססת-קונטקסט

CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production
CARGO היא פלטפורמה שמטפלת בביקורת-מבוססת-קונטקסט של מודלי LLM בייצור. היא חולקת את הביקורת לשלושה סטטוסים: תומך, מנוגד ולא ניתן לביקורת. CARGO-Bench היא סדייה-בדיקה שמאפשרת לבדוק את CARGO.
תקציר מקורי באנגליתarXiv:2609.30471v1 Announce Type: new Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a different entity, so a literal judge penalizes different identifiers, dates, and statuses as errors or hallucinations. We name this failure mode reference-instance divergence (RID). We propose CARGO, a framework that (i) treats retrieved references as procedural exemplars and grounds factual judgments in the live instance's observed context, (ii) assigns each claim a three-way status (supported, contradicted, unverifiable) and penalizes only contradictions, and (iii) gates evaluation by
קרא במקור המקורי