יום שבת, 1 באוגוסט 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

התפתחות לאופקים ארוכים

Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning
▶ צפה כאן — בלי לצאת מהאתר
רוס טיילור וצ'נגקסי טיילור מ-General Reasoning דנים בלמידת חיזוק לאופקים ארוכים. הם מציגים את האתגרים והפתרונות להתפתחות מודלים שיכולים להישאר עקביים לאורך זמן. הם מדגימים זאת עם משימה שבה מודלים קיבלו כסף אמיתי להמר על משחקי כדורגל, אך התקשו לבצע זאת.
תקציר מקורי באנגליתRoss Taylor opens with some history: back in 2022 he worked on Galactica, an early large model for science that briefly crossed the Rubicon on curated high quality data and intermediate reasoning tokens before the reaction overshadowed the work. That obsession, optimizing what happens between the question and the answer, is where this talk on long horizon reinforcement learning picks up. He and Chengxi Taylor of General Reasoning treat long horizon less as a benchmark and more as a mindset: if you want agents that stay coherent over hours, you have to be patient about signal and deliberate about how you spend tokens. The mechanics they walk through are the ones that make long rollouts trainable. Value models reduce variance and help with credit assignment, bootstrapping pulls signal out of
קרא במקור המקורי