יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הסרת תוכן אינה יוצרת עמידות לטמפר

Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models
הסרת תוכן אינה יוצרת עמידות לטמפר במודלים פתוחי משקל. המחקר חושף שהסרת תוכן פוגעני אינה מבטיחה עמידות לטמפר. נמצא כי תכונות נוספות, כגון גאומטריה של גרדיאנט-דסנט, חשובות יותר.
תקציר מקורי באנגליתarXiv:2610.09004v1 Announce Type: new Abstract: Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate is bounded by the fastest reachable parameterization. We apply this principle to weight--data mutual information under training-data filtering and label--representation mutual information under capability removal. Training order can change recovery time at fixed weight--data information, while exact representation-level independence can preserve the entire parameter Jacobian. An explicit construction has both information quanti
קרא במקור המקורי