יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידה מקוונת עם מומחי LLM ממשוב מוגבל

Online Learning with LLM Experts from Limited Feedback
חוקרים פיתחו אלגוריתמים לניתוב פניות למומחי LLM בסביבה מקוונת עם משוב מוגבל. הם השיגו תוצאות טובות בלמידה מהירה של אסטרטגיות ניתוב איכותיות.
תקציר מקורי באנגליתarXiv:2609.05820v1 Announce Type: new Abstract: We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online setting with limited feedback. We formulate it as a bandit problem with $K$ actions that represent experts and $d$ features that encode prompts, over a horizon of $T$ rounds. We propose algorithms that strategically select and observe rewards to minimize regret. In the full-information setting, we achieve a regret of $\tilde{O}(d T / \sqrt{m})$, while in the bandit setting we achieve $\tilde{O}(d T \sqrt{K / m})$, where $m \ll T$ is a budget on feedback. Our experiments show that we efficiently learn high-quality routing strategies across diverse LLMs from limited feedback.
קרא במקור המקורי