יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

MAxBench: תקן חדש לבדיקת ייצוגי רעיונות

MAxBench: A Multinomial Concept Recovery Benchmark
MAxBench הוא תקן חדש לבדיקת ייצוגי רעיונות. הוא נועד לבדוק את יכולת המודלים לשחזור רעיונות רב-קטגוריים. התקן כולל 6 רעיונות שונים ו-4 מודלים שונים. המחקר מצא שהמודלים הטובים ביותר הם אלה שמשתמשים בסביבות ריבועיות.
תקציר מקורי באנגליתarXiv:2609.13072v1 Announce Type: new Abstract: Fine-grained control of language model behaviors (e.g., steering) is among the more actionable outcomes of interpretability research. For binary concepts such as refusal, a single direction in activation space often suffices for steering. However, many concepts are not binary: Animals and Countries contain many subcategories, each with multiple instances. For these concepts, the search space over possible representation geometries is far larger than for binary concepts; it is thus not clear what geometries are most appropriate, nor what methods are most effective at recovering them. In this work, we introduce MAxBench, a geometry-agnostic evaluation framework for multinomial concept representations based on sampling from the recovered concept
קרא במקור המקורי