כתבה
arXiv cs.CL ·
Verifiable Self-Evolution for Open-Ended Dialogue Skills via Future-Feedback Prediction
תקציר מקורי באנגליתarXiv:2607.18973v1 Announce Type: new Abstract: Textual skills provide a lightweight way to improve frozen language-model agents, but their self-evolution normally requires a stable validation signal. Such signals are natural in mathematics or code, where an answer can be checked after it changes, yet are problematic in open-ended dialogue: changing the assistant response also changes the user's next reaction, so a logged reaction cannot directly evaluate a counterfactual response. We propose future-feedback skill evolution, which first redirects self-evolution from prescribing the current answer to predicting whether the observed answer will lead to a positive or negative subsequent user signal. This prediction task is verifiable on fixed logged tuples and therefore supports validation-ga
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית