יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

למידת מדיניות עם צוואר בקבוק שפה

Policy Learning with a Language Bottleneck
חוקרים פיתחו שיטה חדשה ללמידת מדיניות עם צוואר בקבוק שפה, המאפשרת לסוכנים לייצר כללים לשוניים הלוכדים אסטרטגיות ברמה גבוהה. השיטה מוכיחה עצמה בחמישה משימות שונות, כולל משחק אותות וניווט במבוך.
תקציר מקורי באנגליתarXiv:2405.04118v4 Announce Type: replace-cross Abstract: Modern AI systems such as self-driving cars and game-playing agents can achieve superhuman performance, but often lack human-like generalization, interpretability, and inter-operability with human users. Inspired by the rich interactions between language and decision-making in humans, we introduce Policy Learning with a Language Bottleneck (PLLB), a framework enabling AI agents to generate linguistic rules that capture the high-level strategies underlying rewarding behaviors. PLLB alternates between a *rule generation* step guided by language models, and an *update* step where agents learn new policies guided by rules, even when a rule is insufficient to describe an entire complex policy. Across five diverse tasks, including a two-p
קרא במקור המקורי