יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מעבר לנורמה גלובלית: התאמת רגישות לרעילות במודלי שפה ללא אימון מחדש

Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining
חוקרים פיתחו שיטות להתאמת רגישות לרעילות במודלי שפה ללא צורך באימון מחדש. השיטות מועילות להפחתת רעילות בשפה, אך חושפות סתירה בין יעילות ההתאמה לאיכות השפה.
תקציר מקורי באנגליתarXiv:2607.23175v1 Announce Type: new Abstract: Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present the first comparative evaluation of training-free methods for aligning language generation to user-specific toxicity sensitivities across three inference-time intervention stages: pre-decoding (prompt conditioning and rewriting), in-decoding (token, logit, and representation steering), and post-decoding (candidate re-ranking). Evaluated against toxicity sensitivity targets derived from the PRISM dataset, all methods reduce alignment error by 28-47%. However, the results reveal a fundamental trade-off between alignment effectiveness, personalization, and general language quality, showing how toxi
קרא במקור המקורי