יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

JuryFlow: שיפוט רב-סוכנים

JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation
JuryFlow הוא כלי לשיפוט רב-סוכנים שמאפשר הערכה מדויקת יותר של תוכן AI. הוא משתמש בגרף הסכמה כדי לזהות מקרים של אי-הסכמה בין השופטים ולשפר את ההערכה.
תקציר מקורי באנגליתarXiv:2609.40103v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as automated judges for AI-generated content, yet a single judge is unreliable and even a panel of judges leaves a hard residue: when judges disagree, majority voting discards the conflict instead of resolving it. We present JuryFlow, a disagreement-guided, human-in-the-loop multi-agent evaluation framework that treats inter-judge disagreement not as noise to be averaged away, but as a precise, claim-level signal indicating where an evaluation is uncertain. JuryFlow decomposes each candidate response into atomic claims, has a panel of heterogeneous judges assign per-claim verdicts, and builds a disagreement graph whose nodes are scored by verdict entropy and whose edges encode structural
קרא במקור המקורי