וידאו
YT AI Engineer ·
מה בא אחרי RLHF?
What's Next After RLHF? — Diogo Almeida, TypeSafe AI
▶ צפה כאן — בלי לצאת מהאתר
Diogo Almeida, מחבר GPT-4, טוען ש-RLHF יוצרת מודלים שמתמקדים ברצון האדם. הוא מציע כיוונים חדשים לאופטימיזציה, כולל למידת חיזוק עם פרסים מאומתים.
תקציר מקורי באנגליתRLHF made models that are extraordinary at pleasing the human in the loop, and Diogo Almeida, a GPT-4 co author, argues that is exactly the problem. Optimizing for human preference optimizes for engagement and for overpromising, the same pressure that makes a model confidently agree that a fart audio file is a symphony. That produces two camps: one where models act as assistants with a human catching mistakes, where RLHF shines, and one where they operate autonomously with real stakes, where the same instinct to please quietly becomes a liability. So what comes next is not the Claude Code era but a shift in what you optimize. Almeida frames it through Sutton's bitter lesson: the task matters more than the data, and reinforcement learning with verifiable rewards points the model at real aut
קרא במקור המקורי
youtube.com
פתח כתבה מקורית