כתבה
arXiv cs.LG ·
Towards Identifying the Dataset Biases Causing Phantom Transfer
תקציר מקורי באנגליתarXiv:2609.14449v2 Announce Type: replace Abstract: Recent work has shown that a teacher model can transfer a bias to a student through a dataset from which every explicit reference to that bias has been filtered out, and that none of the tested data-level defenses reliably removes or detects such a bias, even when the defender knows what to look for. To shed light on the hidden traces these biases leave, we embed a dataset's completions with Sentence-BERT, subtract the embeddings of clean reference completions, and compare the result to an open vocabulary of candidate topics. This simple signature identifies the topic of the bias with a Matthews correlation coefficient of 0.83 when the attacker's teacher model is known, and 0.46 when it is not. We also observe that different teacher model
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית