יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

VeriHarness: גידול באימות סובייקטיבי למשימות ארוכות-טווח

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
VeriHarness מגדיל את יכולת האימות של סובייקטים למשימות ארוכות-טווח. המאמר מציג פתרון חדשני לבעיה זו, המשתמש בטכנולוגיית LLM כדי לאמת תוצאות של סובייקטים. הפתרון, הנקרא VeriHarness, מסוגל לאמת תוצאות של סובייקטים באופן יעיל ומדויק, ומציע פתרון חדשני לבעיה של אימות סובייקטים.
תקציר מקורי באנגליתarXiv:2610.00972v1 Announce Type: new Abstract: As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreem
קרא במקור המקורי