יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

למידה עצמית בתקופת החוויה

Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?
חוקרים בדקו את היכולת של למידה עצמית בתקופת החוויה, ומצאו שהיסטוריה של הלמידה יכולה להיות נכס או נטל. הם בדקו את השפעת ההיסטוריה על הלמידה, ומצאו שהיא יכולה להרחיב את סולם הפרסים הזמין.
תקציר מקורי באנגליתarXiv:2609.35897v1 Announce Type: new Abstract: The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-driven principles of "Era of Experience". Yet, any self-improving architecture ultimately rests upon its underlying optimization engine: if general intelligence requires learning from grounded interaction, the reinforcement learning (RL) update rule itself must be capable of cumulative adaptation. While algorithm self-discovery has produced Disco103 that surpassed PPO to achieve SOTA benchmark performance -- its internal update machinery remains an uninspected black box. We present the first causal mechanistic audit of a self-discovered RL rule, structured directly around the five pillars of the
קרא במקור המקורי