כתבה
arXiv cs.CL ·
מעבר מעניינים למדיניות: מערכת למידה במקום-זמן יעילה דרך אימיטציה של חקירה של מומחה
From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation
מערכת למידה במקום-זמן יעילה שמשתמשת באימיטציה של חקירה של מומחה. המערכת פותחה על ידי צוות של מדעני דטא ומהנדסים, ומטרתה היא לפתח טכנולוגיה שתאפשר למודלי לשון ללמוד באופן יעיל ומהיר. המערכת משתמשת בשיטת חקירה של מומחה, שבה המודל נתון תפקידים שונים ומחקיר את הבעיה. המערכת נבחנה במספר תחומים, כולל רפואה, טכנולוגיה ומדע. התוצאות היו מוצלחות, והמערכת הציגה יכולות טובות יותר ממודלים אחרים.
תקציר מקורי באנגליתarXiv:2608.16831v2 Announce Type: replace-cross Abstract: Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples. We argue that current prompt and context optimization methods underuse the extensive knowledge and reasoning capabilities of trillion-parameter models. These capabilities can make adaptation more sample-efficient, more compute efficient and at no performance loss when organized around how human experts investigate failures. We formalize Policy Iteration with Human Feedback (PIHF), which makes this implicit procedure explicit for LLM agents to execute, and build its automated implementation, PIHF-MCP. Initialized from clinician feedback on rare-disease diagnosis, PIHF-MCP supplies the expert procedure, testing t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית