יום שלישי, 6 באוקטובר 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

בניית פלטפורמת אימות: מדוע זה קשה יותר ממה שנראה

Why Building an Eval Platform Is Harder Than It Looks — Braintrust
▶ צפה כאן — בלי לצאת מהאתר
הוסיין ניאזמנדי, מהנדס פתרונות בחברת Braintrust, מסביר מדוע מדידת איכות הסוכנים היא יותר מאשר UI על גבי גיליון אלקטרוני. הוא מדבר על שני עמודי האיכות, אימות לפני ייצור וניטור לאחר מכן, ומדוע LLMs לא דטרמיניסטיים דורשים את שניהם. הוא עובר על השלבים שצוותים עוברים: גיליונות, UI מותאם, ניסויים מקבילים עבור מנהלי פרויקטים ומומחי תחום, ולבסוף גלגל שמהפך כישלונות ייצור למקרי בדיקה.
תקציר מקורי באנגליתMost teams start their evals in a spreadsheet. Here's what happens when they try to grow out of it. Hossein Niazmandi, who leads solutions engineering in the West at Braintrust, explains why measuring agent quality is much more than putting a UI on a spreadsheet. He covers the two pillars of agent quality, evals before production and observability after, and why non-deterministic LLMs make both necessary. Then he walks through the stages teams go through: spreadsheets, a custom UI, side-by-side experiments for PMs and domain experts, and finally a flywheel that turns production failures into test cases. At that point the hard part becomes the data: traces that are huge, nested JSON, needing both real-time and long-running queries, which is why Braintrust built BTQL. He closes with agents r
קרא במקור המקורי