כתבה
arXiv cs.CL ·
HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation
תקציר מקורי באנגליתarXiv:2610.03063v1 Announce Type: new Abstract: Large Language Models (LLMs) are prone to generating hallucinated content, which compromises their reliability in knowledge-intensive tasks. To address this challenge without sacrificing creativity, we propose HARPO, a reinforcement learning framework designed to jointly optimize faithfulness and creativity. HARPO incorporates a Hallucination-Aware Generative Reward Model (HA-GRM), trained via verifiable feedback, to assess both faithfulness and writing quality. A Selective Activation Mechanism (SAM) activates writing rewards only for outputs judged hallucination-free by HA-GRM, while a data curriculum progressively shifts training from creative writing to hallucination-centric tasks. On RAGTruth, our Qwen3-4B-based HA-GRM achieves a response
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית