יום שלישי, 6 באוקטובר 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 — Laurie Voss, Arize AI

▶ צפה כאן — בלי לצאת מהאתר
תקציר מקורי באנגליתAn LLM correctness judge rejects all 13 reports from a financial analysis agent. Laurie Voss examines why: the agent used live web research, while the judge graded from its own knowledge without that research context. Supplying the collected sources to a faithfulness evaluator produces a more useful split of six faithful reports and seven unfaithful ones. The notebook builds the agent with the Claude Agent SDK and instruments it with OpenTelemetry and OpenInference to send traces to Arize AX. Reading those traces exposes failed file writes, missing report content, and excessive web searches. A deterministic ticker check adds a cheap first layer of evaluation. Voss builds a custom actionability rubric with explicit criteria, tagged inputs, examples, and binary labels. Each evaluator checks
קרא במקור המקורי