כתבה
arXiv cs.LG ·
GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning
תקציר מקורי באנגליתarXiv:2609.39869v1 Announce Type: cross Abstract: Grammar-constrained generation guarantees syntactic validity, but can substantially degrade semantic quality when the model's preferred outputs are poorly aligned with the imposed grammar. This trade-off is particularly severe when the prompt is underspecified or the model has limited instruction-following ability. Beam search can partially mitigate these failures by exploring multiple valid sequences, but its computational cost grows with beam width, while sequence-level probability is only an imperfect proxy for semantic quality. We introduce GrammarRL, a label-free reinforcement learning method that adapts language models to grammar constraints without requiring annotated data. GrammarRL optimizes the model using two complementary self-s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית