כתבה
arXiv cs.CL ·
יכולת הביקורת משפיעה על דחיית מטרות, לא על כישורי תיקון
Reviewer Capability Governs Rejection Targeting, Not Repair Skill: Evidence from LLM Execute-Review-Revise Pipelines
מחקר חדש בודק את השפעת יכולת הביקורת על דחיית מטרות בצינורות LLM. התוצאות מראות כי ביקורת ברמה בינונית משפרת את הדיוק, אך ביקורת עצמית לא מועילה. המחקר השתמש במודל LLaMA ובמסגרת LangChain.
תקציר מקורי באנגליתarXiv:2609.04270v1 Announce Type: cross Abstract: Multi-agent LLM pipelines increasingly assign roles, including execution and verification, to models of different capability tiers. This is done because running a flagship model at every stage is expensive. Previous literature has established that verification stages are not always beneficial, but holds reviewer capability roughly fixed relative to the executor. We vary it. We replace the reviewer with models spanning a capability range down to one that cannot solve the problems at all, and measure the outcome of every individual rejection. This is done across a constant set of 100 olympiad mathematics problems. A cross-family mid-tier reviewer improves final accuracy by 12 percentage points, from 52 to 64 percent (p = 0.0005), with zero da
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית