יום חמישי, 30 ביולי 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

DeepSWE: בנצ'מרק לקידוד עמיד בזיהום

DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve
▶ צפה כאן — בלי לצאת מהאתר
DeepSWE הוא בנצ'מרק לקידוד עמיד בזיהום, הכולל 113 משימות קידוד שנכתבו מאפס. הבנצ'מרק מודד את היכולת של מודלים כמו GPT ו-Claude לבצע משימות קידוד מורכבות.
תקציר מקורי באנגליתDeepSWE is 113 software engineering tasks written from scratch, not scraped from pull requests, so a model cannot have seen them in training. Each one is a long horizon problem drawn from a real open source repository, authored by engineers who actually maintain that code, with isolated environments and program based verifiers that check observable behavior rather than trusting the model's own account. James Shi's point is that once you remove the contamination the leaderboard stops clustering: strong models pull far ahead and others, Gemini 3.1 Pro among them, fall toward the bottom. The more revealing signal is in how models fail. Some quietly expand a task beyond what was asked, a failure mode DeepSWE scores in its own right, and Claude models did this a good fraction of the time while
קרא במקור המקורי