כתבה
arXiv cs.AI ·
SyRuP: שיפור המענה לפקודות מערכת
SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding
SyRuP הוא כלי לשיפור המענה לפקודות מערכת במודלים גדולים של שפה. הוא משתמש בראש תגמול עם תשומת לב חוצת למידה כדי לייצר ציונים למענה לפקודות. SyRuP מראה תוצאות טובות בהשוואה לשיטות אחרות.
תקציר מקורי באנגליתarXiv:2607.23991v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirements. However, models follow these prompts only implicitly through in-context learning, which can be insufficient for complex or compositional prompts. Existing approaches often require model tuning or response-level reranking, limiting their practicality for lightweight inference-time control. We introduce SyRuP, a decoding-time framework for improving system-prompt adherence while keeping the base LM frozen. SyRuP trains a cross-attention reward head from system-prompt-conditioned preference pairs, treating the system prompt as a separate memory to produce token-level adherence scores. At inference, SyRuP
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית