יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ניהול תצוגה למודלים מבודדים

Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts
ניהול תצוגה למודלים מבודדים הוא בעיה מורכבת. מחקר זה מציג אלגוריתם חדש לניהול תצוגה, Hypentropy Policy Gradient, שמשפר את הביצועים בתנאים אדוורסריים. האלגוריתם מתאים למודלים עם דרגת קריאה נמוכה.
תקציר מקורי באנגליתarXiv:2606.14929v2 Announce Type: replace Abstract: Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models. Despite its practical significance, this problem remains poorly understood under realistic conditions like adversarial queries, bandit feedback, and limited observability of models. We formalize embedding model routing as an adversarial contextual linear bandit with low-rank experts, where contexts are queries, actions are items, and experts are the embedding models working on low-rank latent representation spaces. We first establish that standard regret notions suffer from structural misspecification or statistical intractability, and we identify a log-quadratic policy class that is expressive enough to capture query-depen
קרא במקור המקורי