כתבה
arXiv cs.AI ·
Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs
תקציר מקורי באנגליתarXiv:2607.12985v2 Announce Type: replace Abstract: Aligned language models routinely misreport under non-evidential pressure: they cave to a confident user, yet fail to revise when genuine evidence arrives. We cast this as a failure of internal incentive-compatibility and study the two demands, resist (ignore forbidden pressure) and update (follow licensed evidence), on a Bayesian-witness benchmark with known posteriors, where the same user disagreement is evidence or pressure purely by stated source reliability, removing the evidence/pressure confound by construction. Using interchange interventions rather than probes, we causally localize low-rank report coordinates for answer, confidence, and caveat, establishing causal sufficiency at a late intervention site rather than uniqueness or
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית