יום שבת, 1 באוגוסט 2026 LIVE
AI־INFO

וידאו YT AI Engineer ·

למידת חיזוק ללא תגמולים מאומתים

Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect
▶ צפה כאן — בלי לצאת מהאתר
וויל בראון מ-Priime Intellect דן בלמידת חיזוק ללא תגמולים מאומתים. הוא מציג שיטות לבניית אותות תגמול כאשר אין 'אמת' ידועה. הדבר כולל הגדרת 'שופטים' וזוגות שאלות ותשובות המבוססים על מסמכים אמיתיים.
תקציר מקורי באנגליתReinforcement learning has been easy to sell where the answer is checkable, like math or code, and Will Brown's talk is about everything else. Most valuable tasks have no clean verifier, so Prime Intellect's work is on how you build reward signal when there is no ground truth waiting. He frames RL simply first, a model acting in a harness with tools and skills, getting a reward, and nudging its weights, then asks how you keep climbing once you leave the verifiable island behind. His answer leans on environments as the anchor. You can set up judges, generate question and answer pairs grounded in real documents and repos, and use a reverse direction trick where you hide something, like a bug or a backdoor, so the model can learn to find it again, which conveniently gives you a difficulty dia
קרא במקור המקורי