כתבה
arXiv cs.CL ·
שיפור ביצועי דיבור קטלאני: לימוד מחדש של מודלי שפה גדולים
Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities
במאמר זה, נחקרה שיטת לימוד מחדש (RL) לשיפור יכולות הפשטת טקסט של מודלי שפה גדולים (LLMs) בשפה הקטלאנית. המאמר מציג פונקציית שכר חדשה, מותאמת לסגנון פשטת טקסט מטרה, ומדגים את יעילותה באמצעות GRPO.
תקציר מקורי באנגליתarXiv:2609.04823v1 Announce Type: new Abstract: Although automatic text simplification (ATS) is critical for accessibility, its progress has not matched the rapid evolution of broader natural language processing techniques. This paper investigates the application of reinforcement learning (RL) to improve the quality of ATS for low-resource languages using Large Language Models (LLMs). The paper introduces a novel reward function, designed to guide LLMs toward a targeted simplification style with Group Relative Policy Optimization (GRPO), that combines the SARI metric with specific penalty components. The effectiveness of GRPO with this reward function is motivated and demonstrated by post-training IberianLLM-7B-Instruct on the ASSET dataset. After post-training on the English ASSET, the mo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית