יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

לא כל שאלה היא תקדימה: למידה לתקן תשובות של מומחים בהתקצר

A Query Is Not a Commitment: Learning to Correct Expert Answers in Online Deferral
במאמר זה, חוקרים פיתחו שיטה ללמידה חיצונית שמטרתה לתקן תשובות של מומחים בזמן אמת. השיטה, שנקראת ORUCB, משתמשת בפולינומים על-מומחים ומשותפים כדי לקבוע תקינות של תשובות. החוקרים הציגו תוצאות טובות של השיטה במספר ניסויים.
תקציר מקורי באנגליתarXiv:2610.07084v1 Announce Type: cross Abstract: An inaccurate expert can still provide useful information after correction. We study online learning to defer in which the learner chooses an expert and fixes a correction function before purchasing its answer, then applies that function to the answer received. The difficulty is that observed losses reflect both expert quality and an unfinished correction: early errors can discourage queries that would be valuable after learning. We propose ORUCB, which pools shared and expert-specific polynomial responses. A bound on cumulative response-learning error calibrates confidence-weighted risk regression and exploration, allowing the router to account for this error when deciding which answers to buy. Under bounded residuals and disagreements, a
קרא במקור המקורי