כתבה
arXiv cs.CL ·
OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills
תקציר מקורי באנגליתarXiv:2607.20121v2 Announce Type: replace Abstract: LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra security vulnerabilities, as seemingly harmless skills can contain latent safety risks that only emerge during actual execution. In this work, we conduct a systematic investigation into how well current agent systems recognize and avoid such risks. To support quantitative and qualitative evaluation, we construct OpenSkillRisk, a dedicated safety benchmark containing 263 risky skills collected from public skill marketplaces. We classify these skills into seven categories based on their threat types and pair each skill with a standardized user task and a corresponding sandbox for controlled evalu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית