יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

CARGO: הערכת AI אגנטי בהתאם להקשר

CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production
CARGO הוא כלי להערכת AI אגנטי בהתאם להקשר. הוא משתמש ברעיון של 'reference-instance divergence' כדי להעריך את הביצועים של מודלים כמו LLaMA. CARGO מאפשר הערכה מדויקת יותר של AI אגנטי ביישומים מציאותיים.
תקציר מקורי באנגליתarXiv:2609.30471v1 Announce Type: cross Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a different entity, so a literal judge penalizes different identifiers, dates, and statuses as errors or hallucinations. We name this failure mode reference-instance divergence (RID). We propose CARGO, a framework that (i) treats retrieved references as procedural exemplars and grounds factual judgments in the live instance's observed context, (ii) assigns each claim a three-way status (supported, contradicted, unverifiable) and penalizes only contradictions, and (iii) gates evaluation
קרא במקור המקורי