כתבה
arXiv cs.CL ·
JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety
תקציר מקורי באנגליתarXiv:2607.19913v1 Announce Type: cross Abstract: Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate delayed risks from partial trajectories. Janus synthesizes diverse agent trajectories via multi-agent simulation and learns a shared policy with two coupled tasks: an anticipation task that forecasts safety-relevant futures and an adjudication task that decides safety from both the observed prefix and anticipated future. The two tasks are jointly optimized with CoAA-RL, which rewards forecasts by their utility for downstream safety judgment. The resulting guard model, Vanguard, blocks unsafe actions before execution.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית