יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Mubric: ייצור רוביקה לבדיקת תקינות LLM

Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation
Mubric היא שיטה לייצור רוביקה לבדיקת תקינות LLM, המשתמשת בבדיקת תקינות שינויים. השיטה נבחנה ב-703 תחומים שונים והציגה תוצאות טובות יותר משיטות אחרות.
תקציר מקורי באנגליתarXiv:2609.37322v1 Announce Type: new Abstract: Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubric generation. Mutation testing, a classic software testing methodology, evaluates a test suite by injecting faults into programs and checking whether the tests detect them. We draw an analogy between test suites and rubrics: if a rubric captures an important quality requirement, introducing a corresponding defect into an otherwise high-quality response should reduce its score. Mubric first mines common defects from real pairs of
קרא במקור המקורי