כתבה
arXiv cs.AI ·
סוכנים מדעיים: הערכת פרומפטים מקצועיים
Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks
חוקרים בדקו את היעילות של פרומפטים מקצועיים במשימות מדעיות. הם השתמשו במודל Gemini 3.8 Flash ובמערכת OpenRouter. התוצאות הראו שהפרומפטים המקצועיים לא הביאו לשיפור בדיוק, אך גרמו לעלייה בעלות.
תקציר מקורי באנגליתarXiv:2610.00084v1 Announce Type: new Abstract: Detailed profession-specific system prompts raise token use and estimated cost per response without a consistent accuracy gain. We evaluate Scientific Agents, an open-source corpus of 503 profession-specific AGENTS.md profiles, with Gemini 3.8 Flash via OpenRouter in the Pi agent harness. We compare matched profiles with four controls: a minimal baseline ("You are a helpful assistant"), the profile's opening role sentence, a generic scientific rigor guide, and a profile from an unrelated domain. Across nine text-based science benchmarks (4,531 sampled questions, 100 matched profiles), 4,488 items completed all five conditions after API-error retries, scored with automated, rule-based grading. The average profile-baseline accuracy difference i
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית