יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

מצב הנתיביות של המודלים — NVIDIA, Cognition, OpenRouter

The State of Model Routing — NVIDIA, Cognition, OpenRouter
▶ צפה כאן — בלי לצאת מהאתר
NVIDIA, Cognition ו-OpenRouter דנים בנתיביות מודלים. הם משווים ביצועים ועלות של מודלים שונים, ומציעים פתרונות לשיפור היעילות. הדיון כולל השוואה בין Opus ו-Haiku, והצגת גישות חדשות לניהול מודלים.
תקציר מקורי באנגליתRun terminal bench on Opus and on Haiku and Opus scores about three times better at a tenth of the cost, even though Haiku is far cheaper per token. Alex Atallah's point is that a small model pushed outside its training distribution thrashes, calling tools in loops until it costs more than the expensive model ever would. That inverts the obvious version of model routing, where you send each task to whichever model benchmarks best on it. Walden Yan calls that approach fragile for exactly the reason agents make it worse: a session starts as a question about a codebase, becomes a feature request, then becomes live debugging, and the model you picked at the start is stranded. Cognition's answer keeps a frontier model planning and delegates the implementation, which cut the cost of Fable level
קרא במקור המקורי