כתבה
arXiv cs.AI ·
GrammarRL: פיתוח גרמטיקה-מוגבל באמצעות למידת רב-מודל
GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning
GrammarRL מציע שיטה ללמידת רב-מודל שמתאימה את עצמה לקונסטריינטים גרמטיקליים. השיטה, GrammarRL, משתמשת בשני פרסומים עצמאיים: פרסום ישיר, המדד את הסבירות של הפלט המוגבל, ופרסום הפוך, המדד את היכולת לשחזר את הקלט מהפלט. GrammarRL משפרת את הביצועים של גרמטיקה-מוגבל, ומשמשת כאלטרנטיבה לחיפוש-בקרה.
תקציר מקורי באנגליתarXiv:2609.39869v1 Announce Type: new Abstract: Grammar-constrained generation guarantees syntactic validity, but can substantially degrade semantic quality when the model's preferred outputs are poorly aligned with the imposed grammar. This trade-off is particularly severe when the prompt is underspecified or the model has limited instruction-following ability. Beam search can partially mitigate these failures by exploring multiple valid sequences, but its computational cost grows with beam width, while sequence-level probability is only an imperfect proxy for semantic quality. We introduce GrammarRL, a label-free reinforcement learning method that adapts language models to grammar constraints without requiring annotated data. GrammarRL optimizes the model using two complementary self-sup
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית