כתבה
arXiv cs.AI ·
תיקיות קומפקטיות ל-LLM
Many Preferences, Few Policies: Compact Portfolios for Multi-Objective LLM Alignment
PALM הוא אלגוריתם לזיהוי תיקיות קטנות של LLM ששומרות על ביצועים אופטימליים. הוא משלב רשת מובנית של וקטורי משקל, חיפוש עצלני וקיצוץ. PALM תומך באישוניות מסונכרנת, חקר משקלי פרס במהלך פיתוח מודל ותצורות קומפקטיות בזמן פענוח.
תקציר מקורי באנגליתarXiv:2604.04144v3 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) requires balancing competing objectives such as helpfulness, harmlessness, and conciseness. The appropriate balance varies across users and applications, yet training, evaluating, and deploying many policies across different reward weights is costly. We study how to identify a small portfolio of LLMs that preserves near-optimal performance across all reward weightings. We propose PALM (Portfolio of Aligned LLMs), an algorithm that combines a structured grid of weight vectors, a lazy search that optimizes policies only where needed, and pruning. Given target approximation tolerances, PALM returns a portfolio that provably contains a near-optimal policy for every weight vector, with an explicit up
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית