כתבה
arXiv cs.AI ·
מהו תפקידו של כישור באמת?
What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical Review
ביקורת על מאמרים העוסקים בכישורים וכלים של סוכנים LLM. המחקר בוחן את התוצאות והמסקנות של 100 מאמרים ומציג רשימת בדיקה לשיפור הדיווח.
תקציר מקורי באנגליתarXiv:2609.33153v2 Announce Type: replace-cross Abstract: Reported improvements from tools and reusable skills in large language model agents refer to different comparisons. This critical narrative review examines what these evaluations estimate and which conclusions their designs support. The review checks the roles of one hundred cited papers and extracts focal evaluation designs in detail from thirty-five studies. Targeted readings of thirty-five additional published or accepted studies broaden coverage of tool creation, memory, interactive benchmarks, reliability, and risk. Designs are characterized by treatment contrast, target population, outcome, budget constraint, summary measure, and identification assumptions. Analytic decompositions and counterexamples show that pairing runs on
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית