כתבה
arXiv cs.AI ·
ReSI: Recursive Safety Improvement toward Resistant and Resilient AI
תקציר מקורי באנגליתarXiv:2610.12233v1 Announce Type: cross Abstract: Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment. Models evolve through frequent updates, and their safety alignment requires continual adaptation to each new checkpoint. Meanwhile, with evolving red-teaming methods exposing new vulnerabilities, safety improvement for each checkpoint needs to mitigate exposed vulnerabilities and generalize to risks not yet revealed. Following R$^2$AI, we term these goals resistance to known threats and resilience to unforeseen risks. Recursive self-improvement, in turn, inspires an approach to both goals: safety alignment could likewise adv
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית