כתבה
arXiv cs.AI ·
כשהעולם קובע: התקפי רכבת על דגמי עולם נסתרים לשליטה בשלבי יישום
When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control
המאמר עוסק בהתקפי רכבת על דגמי עולם נסתרים שמשמשים כבסיס לשליטה בשלבי יישום. התקפים אלה יכולים להפעיל שליטה על דגמי עולם נסתרים, כולל גמיני ולאנגרפ.
תקציר מקורי באנגליתarXiv:2609.15781v1 Announce Type: cross Abstract: Pretrained world models, learned simulators that encode an observation into a latent state and predict how it evolves under actions, are beginning to be reused as off-the-shelf dynamics backbones for control, like pretrained encoders and language models are reused today. We show that this reuse opens a supply-chain backdoor: an adversary who controls only a released checkpoint can hijack the downstream controller, even though the victim trains and evaluates entirely on clean data and never sees the trigger. The attack encodes no explicit trigger-to-action rule. Instead, the poisoned model routes trigger-bearing observations into a chosen latent region and reshapes the local dynamics there, so that the victim's own optimization (Dreamer-styl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית