יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

למה 80% אמינות אינה טובה מספיק — Felipe Blanes, Amazon AGI Lab

Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab
▶ צפה כאן — בלי לצאת מהאתר
פליפה בלנס, חבר Amazon AGI Lab, מסביר למה 80% אמינות אינה טובה מספיק לאג'נטים AI. הוא מספר על חוויותיו עם נועה, שירות של Amazon לבניית אג'נטים לדפדפן. הוא מציע לבנות חידושי ערך סביב הלקוח, כדי לצפות ולהתאים לשימושים שלא צפו. הוא מדגיש חשיבות של פתיחות על גבולות המודל, ומספר על שימושים פתאומיים של Nova Act ב-Amazon, Hertz ו-Sola.
תקציר מקורי באנגליתYour agent scores great on benchmarks. Then real customers use it and it breaks. Felipe Blanes from the Amazon AGI Lab shares what he learned working directly with customers of Nova Act, Amazon's service for building browser agents, from research preview to general availability on AWS. He calls the core problem the benchmark illusion: static evals look great until customers do things nobody expected, and then all you have left is hope. His fix is an eval flywheel built around the customer: define success the way the customer does, capture signals (including by talking to customers), diagnose gaps in the model, harness or product, and feed them into decisions. He explains the trust cliff (80% reliability feels like more work; around 92% feels trustworthy), why being open about limits builds
קרא במקור המקורי