וידאו
YT AI Engineer ·
עתיד ה-Evals: מ-LLM כשופט לסוכן כשופט
The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI
▶ צפה כאן — בלי לצאת מהאתר
חברת Arize AI צופה כי ה-Evals יתפתחו מבדיקות סטטיות לשימוש בסוכנים כשופטים. הדבר נובע מהתפתחות הסוכנים, שכעת כוללים תכונות כמו טיפול בכלים ולולאות ארוכות. השופטים החדשים יוכלו לגלות מצבי כשל שלא ניתן לתכנן מראש.
תקציר מקורי באנגליתAcross a dozen eval jobs Arize watches the top teams run, one pattern holds: the eval has to change as fast as the agent it grades. In 2023 an agent was barely more than a prompt; since then reasoning, tool calls, and long multi step loops piled on, and every jump in capability quietly broke the eval that came before. So the evals evolved with them. Deterministic checks catch what you can define up front, LLM as a judge adds the analysis a fixed rule cannot, and the newest step, agent as a judge, hunts for failure modes you would never think to write a check for and can open a pull request to fix what it finds. Aparna Dhinakaran's argument is that this arc, from static checks to an agent grading another agent, is where evals go next. Speaker info: - https://x.com/aparnadhinak - https://www
קרא במקור המקורי
youtube.com
פתח כתבה מקורית