כתבה
arXiv cs.AI ·
FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training
תקציר מקורי באנגליתarXiv:2610.00385v1 Announce Type: cross Abstract: Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and downstream learner utility are distinct objectives. We formalize this selection-to-learning gap and introduce FAER as an auditable full-trajectory replay framework. Its training-free fixed selector is a protocol baseline; FAER-UTILITY is the learner-aware selector fitted on disjoint calibration blocks. The normalized gradient alignment is reported as a baseline, while a disposable optimizer-aware virtual update supplies a magnitude-aware utility surface. The audit contract freezes observed fields and replay traces before evaluation labels are joined. On GSM8K with Qwen2.5-1.5B-Instruct, the matc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית