כתבה
arXiv cs.AI ·
מחקר אמפירי בנושא יעילות כלי סקיל
An Empirical Study of Agent Skills' Downstream Utility
במחקר זה נבחנה יעילות כלי סקיל בשימוש במודלי LLM. המחקר נערך על 87 תרגילים וכלל ניתוח של תוכן, ריצות ותוצאות. התוצאות הראו שכלי סקיל יכולים להגביר או להפחית את תוצאות התרגילים, ושרירות תוכן והקמה יכולים להשפיע על יעילותם.
תקציר מקורי באנגליתarXiv:2610.08875v2 Announce Type: replace Abstract: Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Skill organization. We conduct an empirical study on 87 SkillsBench tasks, defining downstream utility as the pass-rate difference from No-Skill on the same tasks under the same model--harness configuration. We compare the same Skills across nine configurations, then examine alternative published Skills and organizations of fixed Skill sets under three selected configurations. We retrieve marketplace candidates from a curated
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית