יום שבת, 1 באוגוסט 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

ללמוד בעבודה: עתיד האימון אחרי ההפלטה

Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute
▶ צפה כאן — בלי לצאת מהאתר
חברת Applied Compute מאמנת מודלים מותאמים עם למידת חיזוק. המודלים משתלבים בתשתיות קיימות ומתאמנים על סמך אינטראקציות מציאותיות, תוך שימוש בלולאת אימון דומה לזו המשמשת בלמידת חיזוק.
תקציר מקורי באנגליתThe next step after a model ships is teaching it to keep learning on the job, and Raymond Feng lays out how Applied Compute trains custom models with reinforcement learning that plug into whatever harness an enterprise already runs. The setup is an orchestrator that fans interactions out to inference engines, collects the graded rollouts, and feeds a training engine that updates the weights, the same GRPO style loop used for RL today, but pointed at real multi turn, long horizon work rather than toy question and answer pairs. The promise is a model you deploy once that adapts to a specific company's tasks. The hard parts are all about the environment. Feng is candid about reward hacking, where a model learns to time out a tool or exploit a scoring gap instead of doing the task, and about t
קרא במקור המקורי