כתבה
arXiv cs.AI ·
מעבר מעניינים למדיניות: מערכת למידה במקום-זמן יעילה דרך הדמיה של החקירה המומחה
From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation
מערכת למידה במקום-זמן יעילה שמשתמשת בהדמיה של החקירה המומחה. המערכת, שפותחה על ידי צוות של קלוד, מאפשרת למודלי LLM ללמוד תכונות חדשות באופן יעיל יותר. המערכת נבחנה במספר תחומים, כולל ריאליות ומתמטיקה.
תקציר מקורי באנגליתarXiv:2608.16831v2 Announce Type: replace Abstract: Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples. We argue that current prompt and context optimization methods underuse the extensive knowledge and reasoning capabilities of trillion-parameter models. These capabilities can make adaptation more sample-efficient, more compute efficient and at no performance loss when organized around how human experts investigate failures. We formalize Policy Iteration with Human Feedback (PIHF), which makes this implicit procedure explicit for LLM agents to execute, and build its automated implementation, PIHF-MCP. Initialized from clinician feedback on rare-disease diagnosis, PIHF-MCP supplies the expert procedure, testing tools,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית