כתבה
Import AI ·
AI מבצעים משימות תכנות למשך שבוע
Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker
חברות Epoch ו-METR השיקו את MirrorCode, בנצ'מרק לבדיקת יכולתם של מודלי AI לבצע משימות תכנות ארוכות טווח. הניסויים הראו כי מודלים כמו Claude Opus 4.7 ו-GPT-5.5 מסוגלים לבצע משימות מורכבות בזמן קצר יחסית.
תקציר מקורי באנגליתWelcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Epoch and METR release MirrorCode, a benchmark for seeing how well AI systems can do long-horizon programming tasks: …AI systems can’t solve the hardest tasks yet (good!)… Epoch and METR have released MirrorCode, a benchmark meant to see how well AI systems can do tasks that take humans a long time to do. The benchmark was first announced in April (Import AI #453) and has now been fleshed out and released with additional tests. The findings are already very striking; Opus 4.7 solved a task in 14 hours for $251 in inference cost which METR and Epoch believe would take a human 2-17 weeks to do. “We also fou
קרא במקור המקורי
jack-clark.net
פתח כתבה מקורית