כתבה
arXiv cs.LG ·
An Empirical Study of Agent Skills' Downstream Utility
תקציר מקורי באנגליתarXiv:2610.08875v2 Announce Type: replace-cross Abstract: Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Skill organization. We conduct an empirical study on 87 SkillsBench tasks, defining downstream utility as the pass-rate difference from No-Skill on the same tasks under the same model--harness configuration. We compare the same Skills across nine configurations, then examine alternative published Skills and organizations of fixed Skill sets under three selected configurations. We retrieve marketplace candidates from a cu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית