יום חמישי, 30 ביולי 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

בניית Evals שדואגות למציאות

Build Evals That Actually Matter - Nick Ung, Lyft
▶ צפה כאן — בלי לצאת מהאתר
חברת Lyft פיתחה סימולטור משתמשים אדוורסרי לבדיקות אופליין. הסימולטור מבוסס LLM מותאם היטב, המדמה משתמשים מבולבלים, כועסים ואדוורסריים. הגישה הזאת מאפשרת לחברה לזהות בעיות בשלב מוקדם.
תקציר מקורי באנגליתYour agent passes offline evals at 90%. You ship. Production immediately finds failure modes your eval never saw. Sound familiar? The culprit is almost always the same: the "customer" in your offline eval is an off-the-shelf LLM that sounds nothing like your real users, and your synthetic test set doesn't capture how messy, angry, or off-topic real conversations get. Your eval was too easy. At Lyft, our customer-care agents resolve roughly a third of all customer issues — millions of conversations a month. To trust them at that scale, we built an adversarial user simulator: a fine-tuned LLM trained on real Lyft rider and driver transcripts that can role-play frustrated, confused, and adversarial users with the same distribution as production. It found regressions our synthetic dataset miss
קרא במקור המקורי