יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אפשרות להסרת תוצאות שאינן רצויות במודלי טקסט-תמונה

Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models
אפשרות להסרת תוצאות שאינן רצויות במודלי טקסט-תמונה. ניתן להסיר תוצאות שאינן רצויות במודלי טקסט-תמונה, כולל תמונות NSFW.
תקציר מקורי באנגליתarXiv:2610.10859v1 Announce Type: cross Abstract: Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference. However, their ability to generate harmful or undesired content poses significant safety risks. Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives. Crucially, these objectives implicitly rely on multi-step denoising dynamics, an assumption that breaks down for few-step distilled (FSD) models, resulting in ineffective forgetting. Furthermore, performing unlearning on the non-distilled base model and subsequently re-distilling it to obtain an unlearned FSD model incurs substantial computational and time overhead, making it impractical in man
קרא במקור המקורי