כתבה
arXiv cs.AI ·
Generalising from Self-Produced Data: Model Training Beyond Human Constraints
תקציר מקורי באנגליתarXiv:2504.04711v2 Announce Type: replace Abstract: Current large language models (LLMs) are constrained by human-derived training data and limited by a single level of abstraction that impedes definitive truth judgments. This paper introduces a novel framework in which AI models autonomously generate and validate new knowledge through direct interaction with their environment. Central to this approach is an unbounded, ungamable numeric reward - such as annexed disk space or follower count - that guides learning without requiring human benchmarks. AI agents iteratively generate strategies and executable code to maximize this metric, with successful outcomes forming the basis for self-retraining and incremental generalisation. To mitigate model collapse and the warm start problem, the frame
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית