יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

TrojanWorld: התקפת Backdoor על סוכנים

TrojanWorld: Backdooring World-Model Agents via Imagination Steering
TrojanWorld היא שיטה להתקנת Backdoor בסוכנים המבוססים על מודלי עולם. היא מאפשרת השתלת התנהגות מזיקה באמצעות הדמייה פנימית. המחקר הראה כי השיטה יכולה להשפיע על סוכנים שונים, כולל TD-MPC2, DreamerV3 ו-R2-Dreamer.
תקציר מקורי באנגליתarXiv:2609.07051v1 Announce Type: new Abstract: World models increasingly serve as the predictive core of model-based reinforcement learning agents, enabling them to simulate future dynamics and reason over imagined trajectories before acting. Their substantial training demands make pretrained world models attractive for distribution and reuse, exposing downstream systems to model supply chain threats. Backdoor attacks offer a targeted and stealthy means of exploiting such supply chains, yet their threat to interactive world-model agents remains largely unexplored. To fill this gap, we present TrojanWorld, a backdoor framework for world-model agents that induces attacker-specified behavior by steering internal imagination. A physical object placed in the scene acts as the trigger, enabling
קרא במקור המקורי