וידאו
YT AI Engineer ·
Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase
▶ צפה כאן — בלי לצאת מהאתר
תקציר מקורי באנגליתThe official judge said the agent succeeded 74% of the time. A better verifier said 38%. Miguel González Fernández, tech lead for Browserbase's agent platform, and Corby Rosset, researcher at Microsoft Research, present their research on verifiers for computer-use and web agents. Deterministic evals stopped scaling as agents improved, and the LLM judges bundled with popular web benchmarks turned out to be confidently wrong. Train against them and you get a more confident liar, not a better agent. They explain how the Universal Verifier works: task-specific rubrics, ranking the most relevant screenshots as evidence for each criterion, isolating errors, catching hallucinations and separating controllable from uncontrollable failures. It cut false positives from about half to near zero and ag
קרא במקור המקורי
youtube.com
פתח כתבה מקורית