יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

כאשר נתונים סינתטיים פוגעים

When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents
מחקר חדש מראה כי נתונים סינתטיים יכולים לפגוע בביצועים של סוכנים LLM. המחקר בדק את השימוש בנתונים סינתטיים לצורך אימון סוכנים LLM ומצא כי הדבר יכול לגרום לשכחה קטסטרופלית. החוקרים הצליחו למצוא פתרונות לבעיה זו באמצעות שיטות שונות.
תקציר מקורי באנגליתarXiv:2609.10750v1 Announce Type: cross Abstract: LLM agents increasingly rely on external skills retrieved at runtime, making skill selection from large repositories a critical challenge. We present a production skill router over 34,396 skills and a large-scale study of skill retrieval using limited real supervision and synthetic data. We found that the synthetic-data fine-tuning improves in-distribution retrieval but it causes catastrophic forgetting on real and out-of-distribution (OOD) data. We evaluate several forgetting mitigation fine-tuning approaches inspired by continual learning, including embedding-anchor regularization, Learning without Forgetting (LwF), Elastic Weight Consolidation (EWC), and L2-initialization. The results show that these approaches not only retain the perfor
קרא במקור המקורי