יום חמישי, 30 ביולי 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

מעבר מעקבות אג'נטים לסימולציות אג'נטים

From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI
▶ צפה כאן — בלי לצאת מהאתר
רוסטם פייזחאנוב מ-Snorkel AI מציג גישה לבניית סימולציות אג'נטים מעקבות ייצור. המטרה היא ליצור סביבת בדיקות פרטית המדמה את התנאים המציאותיים, ולאפשר השוואה בין מודלים שונים. הסימולציה כוללת כלים, מערכות ומדרגים, ומאפשרת לבדוק את ביצועי האג'נטים בתנאים מבוקרים.
תקציר מקורי באנגליתTake a real production trace, rebuild the database state, tools, and files the agent touched, and you have a task any model can replay under identical conditions. That reconstruction is the move at the center of this talk. Public benchmarks like WebArena hand you a single success rate on someone else's tasks, but what you actually care about is cost per solved task, latency, and whether the agent followed your policies. So you build a private benchmark from your own traces, wire in the same skills, tools, and evaluators the agent sees in production, and compare models apples to apples on the environment that matters to you. The environments are multistep and long horizon, so a verifier reads the final state while an LLM judge checks whether the agent followed policy, and a run can stop ear
קרא במקור המקורי