כתבה
arXiv cs.CL ·
הקיצור ההיררכי של תצוגות בנקודת-ביקורת של מודלי תצוגה-שפה
Hierarchical Compression of Vision-Language Model Benchmarks
פרויקט חדש מציע פרקטיקה להקטנת עלויות בבדיקת מודלי תצוגה-שפה. הפרויקט, PRIMEBench, פועל בארבעה שלבים: טיהור נתונים, זיהוי נציגי קטגוריה, קצירת פריטים עם Vision-Aware Variance וקצירת קטגוריות. הפרויקט יכול להקטין את עלויות הבדיקה בכ-97% ולשמור על דירוגי המודלים.
תקציר מקורי באנגליתarXiv:2609.37515v1 Announce Type: cross Abstract: Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that preserve model rankings at a fraction of the cost are well studied for language models, but for VLMs the question remains under-explored. We present PRIMEBench (Pruning Redundant Items for Multimodal Evaluation), a vision-aware hierarchical benchmark compression framework that substantially reduces evaluation cost while preserving model rankings. This hierarchical framework operates in four stages: data cleaning to remove items answerable without the image and all-correct items, category representative selection to p
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית