יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מעבר לניקוד מצטבר: הנחות תקינות התנהגותית

Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods
חוקרים מציגים מסגרת חדשה להערכת שיטות הערכה אוטומטיות המבוססות על התייחסות. המחקר בוחן מגוון רחב של מודלים וכלים, כולל LLM, ומנתח את התנהגותם.
תקציר מקורי באנגליתarXiv:2609.05289v1 Announce Type: new Abstract: Automated reference-based evaluation methods play a critical role in assessing natural language generation systems. Existing meta-evaluation primarily measures agreement with human judgments or benchmark labels, providing limited insight into evaluator behavior under controlled conditions. We introduce behavioral correctness assumptions, a complementary framework for evaluating reference-based automatic evaluation methods. We define a taxonomy of correctness-preserving and correctness-altering assumptions and operationalize them through controlled response transformations that specify expected scoring behaviors. We evaluate diverse lexical, character-level, semantic, LLM-based, and hybrid evaluators and analyze their assumption-level behavior
קרא במקור המקורי