כתבה
arXiv cs.AI ·
Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning
תקציר מקורי באנגליתarXiv:2607.23394v1 Announce Type: new Abstract: Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferences that generalize broadly. Standard defenses, such as data filtering, mixing in harmless data, and regularization, attenuate these effects but do not eliminate them. We instead pursue robustness through redundancy: collecting multiple datasets from different sources and only learning what is common between them. Thus, if only a subset of sources are malicious, the misbehavior will be blocked. In order to implement this defense strategy, we fine-tune a separate reference model on each source's dataset and aggregate their next-token distributions at decoding time. W
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית