יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

ActTraitBench: גיאוגרפיה של הפער בין ידע להחלטה במודלי שפה גדולים דרך התקשורת האנושית

ActTraitBench: Quantifying the Knowledge-Decision Gap in Large Language Models via Human-Grounded Behavioral Validation
ActTraitBench מדד את יציבות הדמות במודלי שפה גדולים. המאמר מציג פלטפורמה לבדיקת תכונות דמותיות ב-LLMs, ומציע פתרון לפער הידע-החלטה. ניתן למצוא את הקוד והמשאבים ב-https://github.com/Selina233/ActTraitBench.
תקציר מקורי באנגליתarXiv:2605.29791v2 Announce Type: replace Abstract: While Large Language Models (LLMs) can convincingly simulate personas in explicit self-reports, they often deviate in implicit behavioral decisions, revealing a substantial Knowledge-Decision Gap ($G_{\mathrm{KD}}$). Existing benchmarks struggle to measure this discrepancy due to limited construct validity, multidimensional entanglement, and distributional biases in LLM-based evaluation. To address these issues, we propose ActTraitBench, a human-grounded evaluation framework for measuring personality consistency in LLMs. Grounded in empirical human data, ActTraitBench establishes one-to-one mappings between psychometric facets and behavioral paradigms and applies Distributional Calibration via Quantile Mapping to reduce distributional mis
קרא במקור המקורי