כתבה
arXiv cs.CL ·
לימוד לציין ממבט משוער של המאזין
Learning to Refer from Estimated Listener Gaze
חוקרים הצליחו לשפר מודלים של שפה-ראייה על ידי שימוש במבט משוער של המאזין. המודלים החדשים מייצרים ביטויים מתאימים יותר למציאות, ומקטינים את אורך הביטויים מ-15.4 מילים ל-4.0 מילים, תוך הגדלת הצלחת הייחוס מ-75.2% ל-80.0%.
תקציר מקורי באנגליתarXiv:2609.14207v1 Announce Type: new Abstract: We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental listener comprehension, in the form of gaze scanpaths, into learning signals. During training, referring expressions are sampled from the speaker policy being optimized, conditioned on images and target referents; then, a neural listener estimating human gaze behavior maps from images and sampled referring expressions to scanpaths, each represented by a sequence of fixations, with each fixation corresponding to a word in the referring expression. We experiment with several approaches to convert fixation sequences and target referents into token- and sequence-level rewards, which are used to opt
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית