כתבה
arXiv cs.CL ·
בנק אפקטיביות: בדיקת עומס של חבילות חיבור למודלי שפה לתגובה על רקע רחב
LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning
במאמר זה נוצרה בדיקה חדשה לבדיקת חבילות חיבור למודלי שפה לתגובה על רקע רחב. הבדיקה כוללת תרגילים שונים שדורשים טקטיקות שונות לחיפוש ולתגובה. המחברים ניתחו את יעילותן של חבילות חיבור שונות ומצאו שאותו מודל יכול להיות יעיל באופן שונה תחת חבילות שונות.
תקציר מקורי באנגליתarXiv:2609.38137v1 Announce Type: new Abstract: Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this paper, we introduce a benchmark for evaluating both the effectiveness and efficiency of long-context harnesses. Our tasks require diverse retrieval strategies, including lexical search and semantic matching, together with strategic and adaptive reasoning over global and local context. Much of the context is semantically relevant but only a small subset is useful at each step, creating both a challenging search problem and different accuracy--cost
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית