כתבה
arXiv cs.AI ·
החלטה לפני צפייה: למידת זיכרונות שראוי להציג בפיקסלים
Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels
PixelTriage הוא אלגוריתם שלומד לזהות אילו זיכרונות שיש להציג בפיקסלים. הוא משתמש בדגם קטן שקורא את הדיאלוג, הערות קצרות ותמונות מוקטנות של זיכרונות שנאספו. האלגוריתם מוכשר על פי פרקי זיכרון סינתטיים שסומנו על ידי דגם 27B קפוא. הוא משתמש ב-11-23% מהפיקסלים ללא אובדן משמעותי בדיוק.
תקציר מקורי באנגליתarXiv:2610.07984v1 Announce Type: cross Abstract: Multimodal assistants answer questions from long-term memories that contain images. After retrieval, each retrieved image reaches the answering model either as pixels, at about a thousand visual tokens per image, or as a stored text proxy that often misses the detail the question asks about. We find that the benefit of pixels usually comes from one or two retrieved memories, and that it can be predicted before the answering model runs, without reading any full-resolution image. In PixelTriage, a plug-in placed after retrieval, a small model that does not generate text reads the dialogue, a short note and a thumbnail of each retrieved memory and predicts how much its pixels would add. It is trained on synthetic memory episodes labeled by a f
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית