כתבה
arXiv cs.AI ·
מדדי זמינות נמוכים, חסרי ראייה: העלות האמינות של סדרות סיירה זולות
Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades
במאמר זה, המחברים מדדים את העלות האמינות של סדרות סיירה זולות שמשתמשות במודלי LLM זולים ומעבירות את האישורים הקשים למודלים חזקים יותר. התוצאות היו חסרי ראייה ולא אמינים.
תקציר מקורי באנגליתarXiv:2609.01345v2 Announce Type: replace Abstract: Inference cascades cut cost by answering most queries with a cheap model and escalating a hard tail to a frontier model that acts as verifier. A natural extension closes the loop: fine-tune the cheap student on the verifier's rejections so the escalation rate, and cost, fall each round. We measure this loop on real LLMs and report four findings. First, the verifier's blind spot, the fraction of the student's wrong answers it accepts, is large and moves adversarially: it grows with student capability ($\beta$ from 0.12 to 0.55 as the student scales 0.5B to 32B) and shrinks with verifier capability, so it is worst in the cheap-student, cheap-verifier regime cascades exist to create. Second, buying it away returns the saving: a frontier veri
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית