וידאו
YT AI Engineer ·
בנצ'מרקים: הטוב, הרע והמכוער
Benchmarks: The Good, the Bad, and the Ugly — Ali Khial, G2i
▶ צפה כאן — בלי לצאת מהאתר
אלי ח'יאל חשף בעיות בבנצ'מרקים פופולריים. הוא מציג דוגמאות שבהן הוראות אינן ברורות, ותוצאות מוטות. ח'יאל מציע עקרונות לבניית בנצ'מרקים אמינים.
תקציר מקורי באנגליתAli Khial took three of the best engineers at G2i, pointed them at popular coding benchmarks, and hit a wall of tasks that were either too ambiguous to grade or quietly broken. That experience is the spine of this talk: a benchmark starts as a spec, solutions get verified and graded, and the results rank models, but only if the harness is actually creating a fair test rather than an unfair one. He shows real examples where an instruction is so vague that a correct patch gets rejected, or a test checks something as arbitrary as how a variable is named, and notes that a meaningful share of tasks he examined had genuinely good answers marked wrong. The danger is that models are increasingly good at gaming exactly this, hunting down the test and satisfying it rather than solving the problem, w
קרא במקור המקורי
youtube.com
פתח כתבה מקורית