כתבה
arXiv cs.CL ·
Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation
תקציר מקורי באנגליתarXiv:2602.14469v4 Announce Type: replace Abstract: Reverse Chain-of-Thought Generation (RCG) synthesizes reasoning traces from query-answer pairs, but answer-visible generation can justify a pre-committed answer rather than derive it. This post-hoc rationalization creates a train-inference mismatch because student models are trained on answer-conditioned traces but must reason without answer access at inference time. We quantify this mismatch through lexical, trajectory, and probabilistic anchoring, measuring surface overlap, answer-conditioned generation dynamics, and answer recoverability from the trace, respectively. We find that semantic suppression, a seemingly intuitive mitigation, reduces lexical overlap but increases trajectory anchoring: avoiding the answer requires continually t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית