יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אישור עצמי באגנטים רזיונינג: חידוש בלמידת מודלים

Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning
אגנטים רזיונינג: חידוש בלמידת מודלים
תקציר מקורי באנגליתarXiv:2609.08025v1 Announce Type: new Abstract: Reasoning agents increasingly rely on external tools such as web search to answer complex queries. Reinforcement learning (RL) finetuning algorithms such as GRPO have improved long-form reasoning in text-only language models, particularly for coding and mathematics. Reliable tool use in multimodal agents, however, remains challenging because models must interpret text and images while integrating noisy retrieved evidence, often under sparse outcome-level supervision without explicit verification signals. We present Self-Verification via Reinforcement Learning (SVRL), an RL-only finetuning framework that trains multimodal agents to verify and filter retrieved evidence within their own reasoning traces, reducing reliance on external verifiers a
קרא במקור המקורי