כתבה
arXiv cs.AI ·
Palette: פלטין - פלטפורמה מודולרית לבקרת סימטריה והתאמה למודלי LLM
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs
פלטין היא פלטפורמה מודולרית שמאפשרת התאמה ובקרת סימטריה במודלי LLM, כולל מודלי Claude, Gemini ו-GPT-5. הפלטפורמה מאפשרת התאמה ובקרת סימטריה בצורה מודולרית ומאפשרת התאמה ובקרת סימטריה בצורה יעילה. הפלטפורמה נועדה לספק פתרון לבעיית הבקרת סימטריה במודלי LLM.
תקציר מקורי באנגליתarXiv:2605.24154v2 Announce Type: replace Abstract: Current safety alignment of foundation models largely follows a \emph{one-size-fits-all} paradigm, applying the same refusal policy across users and contexts. As a result, models may refuse requests that are unsafe for general users but legitimate for authorized professionals, limiting helpfulness in specialized professional settings. Existing approaches either require costly realignment or rely on inference-time steering that suffers from imprecise control and added latency. To this end, we propose \textsc{Palette}, a modular, controllable, and efficient framework that selectively relaxes refusal behavior on authorized target domains while preserving standard safety elsewhere. Our method identifies a refusal direction via multi-objective
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית