כתבה
arXiv cs.AI ·
VeriFine: פלטפורמת ניטור לשיפור עצמי בתפיסה גופנית
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
VeriFine היא פלטפורמת ניטור שמספקת שיפור עצמי בתפיסה גופנית. היא משתמשת בשיטת Policy Improvement Loop ו-Judge Improvement Loop לשיפור יכולת המדיניות והשופט. ניתן לראות את התוצאות במשימות נהיגה וניווט רובוטי.
תקציר מקורי באנגליתarXiv:2610.08761v1 Announce Type: new Abstract: Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and ver
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית