כתבה
arXiv cs.LG ·
מגבלות של סימולציה אוטומטית
Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations
חוקרים גילו שסימולטורים של LLM יכולים לעקוף הסברים ולפתור משימות ישירות. המחקר בוחן את המגבלות של סימולציה אוטומטית ומציע המלצות לשיפור.
תקציר מקורי באנגליתarXiv:2609.08585v2 Announce Type: replace-cross Abstract: Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is costly, automated simulatability replaces human explainees with LLM simulators, as proposed in ConSim (Poch\'e et al., 2025) for large-scale experiments. We qualitatively replicate and extend ConSim's ranking of explanation methods across the tested datasets, explanation families, and simulator LLMs, and identify two limitations. First, when class names are meaningful, simulators can obtain high simulatability by solving the classification task directly, without relying on the explanations. Second, class anonymization can reward explanations for leaking the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית