יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

K-Bench: מדידת ביצועי מודלים בבקשות מדעיות אמיתיות

K-Bench: measuring model performance on real scientific agent requests
K-Bench הוא כלי למדידת ביצועי מודלים בבקשות מדעיות אמיתיות. הוא בוחן את היכולת של מודלים כמו GPT-5.6-sol ו-Claude-opus-5 לענות על בקשות מדעיות מורכבות. התוצאות מראות שאף מודל לא עובר את הרף הנדרש.
תקציר מקורי באנגליתarXiv:2608.21601v2 Announce Type: replace-cross Abstract: Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure. Real scientific requests arrive differently. They are underspecified, they carry attachments, and they lack ground truth. We report K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs. Three blinded language-model judges scored every run against an eight-dimension rubric. On a rubric whose 8-anchor is defined as work a domain scientist would accept with minor edits, no model clear
קרא במקור המקורי