יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

וידאו YT IndyDevDan ·

דירוג מהנדסי סוכנים: איך אני מדרג Astra, Fable 5.1 ו-Open-Weights

Agentic Engineering Benchmarks: How I RANK Astra, Fable 5.1, and Open-Weights
▶ צפה כאן — בלי לצאת מהאתר
דירוג מהנדסי סוכנים מושווה את GPT-6 Astra, Claude Fable 5.1 והשדה הפתוח של Open-Weights. הדירוג מבוסס על חמישה בנצ'מרקים: Terminal-Bench v4.0, APEX Agents Leaderboard, AutomationBench, AA-Omniscience ו-DeepSWE v1.1. הדירוג בוחן את הביצועים, העלות והמהירות של כל מודל.
תקציר מקורי באנגליתThe Artificial Analysis Index is lying to you. 😲 Not on purpose, but an index is a proxy of a proxy, and by the time ten benchmarks get mashed into one number, the only thing that actually matters (which model to run for YOUR work) is gone. 🎥 VIDEO REFERENCES • Claude Fable & Mythos 5.1: https://www.anthropic.com/claude-fable-and-mythos-5-1 • GPT-6 Astra: https://openai.com/index/gpt-6-astra/ 📊 MY TOP 5 BENCHMARKS • Terminal-Bench v4.0: https://artificialanalysis.ai/evaluations/terminalbench-v4-0 • APEX Agents Leaderboard: https://www.mercor.com/apex/apex-agents-leaderboard/ • AutomationBench: https://artificialanalysis.ai/evaluations/automationbench-aa • AA-Omniscience: https://artificialanalysis.ai/evaluations/omniscience • DeepSWE v1.1: https://deepswe.datacurve.ai/blog/deepswe-v1-1
קרא במקור המקורי