כתבה
arXiv cs.CL ·
HARDE: Optimizing Agent Harnesses for Runtime Risk Detection and Execution Control
תקציר מקורי באנגליתarXiv:2609.38291v1 Announce Type: cross Abstract: Large language model (LLM) agents are vulnerable to safety risks such as injected malicious instructions or misleading information, motivating runtime defenses that prevent unsafe action in execution across diverse risks while preserving benign-task utility. Existing system-level defenses either focus on risk detection rather than timely prevention or rely on predefined rules with limited flexibility across diverse risks. We propose a risk-aware harness that integrates LLM-based monitoring for flexible risk detection and structures monitor-guided execution around three core modules: trigger, monitor, and feedback, enabling targeted safety interventions while limiting disruption to benign task execution. To adapt the harness to different ris
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית