כתבה
arXiv cs.CL ·
קטן ומשרת: תאימות-מרחבית-אחרי-אימון לראיה-נמוכה
Small yet Assistive: Spatially-Aware Post-Training for Low Vision
מודל קצר למשתמשים עיוורים ועיוורי ראייה שמספק פרטים מרחביים והתרשמות מסכנות.
תקציר מקורי באנגליתarXiv:2609.28757v1 Announce Type: cross Abstract: An estimated 1 billion people worldwide live with vision impairment, yet current vision-language models (VLMs) produce descriptions too vague for safe navigation by blind and low-vision (BLV) users. Large VLMs can generate high-quality audio-description-compliant narrations but cannot run on mobile devices; small VLMs offer competitive latency but lack spatial detail, directional cues, and hazard awareness for navigational assistance. We present Smol-VL-BLV, a compact VLM for blind and low-vision users that closes this gap using a 500M decoder transformer model and two post-training mechanisms: (1) teacher-student distillation and (2) Group Relative Policy Optimization (GRPO) with a composite BLV reward targeting directional language, metri
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית