כתבה
arXiv cs.AI ·
גילוי וחיסון תגמולים מנוצלים במודלים שפה עצמאיים
Harness-agnostic detection and immunization of reward hacking in self-evolving language models
חוקרים פיתחו שיטה לגילוי וחיסון תגמולים מנוצלים במודלים שפה עצמאיים. השיטה, HackProbe, משתמשת בליבה קבועה ושכבה חדשה לאיתור ומניעת התנהגויות לא רצויות. היא הוכחה יעילה במניעת תגמולים מנוצלים ושיפרה את היכולת הכללית של המודלים.
תקציר מקורי באנגליתarXiv:2609.04665v1 Announce Type: new Abstract: Self-evolving language models improve by proposing candidate updates and keeping whatever raises a visible score. When that score is an imperfect proxy for the capability one actually wants, sustained selection widens the gap between the two. This is reward hacking. We introduce HackProbe, a monitor that attaches to an arbitrary self-evolving loop through two black-box hooks, with no access to weights or activations. It keeps a secret, distribution-fixed comparison core, whose frozen distribution makes its capability proxy comparable across generations, alongside a rotated fresh layer that hardens the bank against co-adaptation. Four tests built on that proxy cover the level gap, a scale-aligned divergence with online change-point detection,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית