וידאו
YT AI Engineer ·
Evals in AI: A Deep Dive — Tejas Kumar, IBM
▶ צפה כאן — בלי לצאת מהאתר
תקציר מקורי באנגליתA customer support answer passes a test because it contains the word “cannot,” even though it approves a return outside the store's policy. Tejas Kumar builds that failure live to show where ordinary assertions and fuzzy matching stop being useful. The example grows into an LLM judge, then a dataset of customer scenarios, agent responses, and human verdicts. Along the way, the judge starts favoring an answer generated by its own model family. Adding the actual return policy improves its decisions, but missing details and a loyal customer's appeal still expose gaps. Kumar keeps examining disagreements, supplying context, and eventually changing the model rather than treating a green test as proof that the system works. The workshop connects those experiments to an evaluation process that ca
קרא במקור המקורי
youtube.com
פתח כתבה מקורית