יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Towards Identifying the Dataset Biases Causing Phantom Transfer

תקציר מקורי באנגליתarXiv:2609.14449v2 Announce Type: replace Abstract: Recent work has shown that a teacher model can transfer a bias to a student through a dataset from which every explicit reference to that bias has been filtered out, and that none of the tested data-level defenses reliably removes or detects such a bias, even when the defender knows what to look for. To shed light on the hidden traces these biases leave, we embed a dataset's completions with Sentence-BERT, subtract the embeddings of clean reference completions, and compare the result to an open vocabulary of candidate topics. This simple signature identifies the topic of the bias with a Matthews correlation coefficient of 0.83 when the attacker's teacher model is known, and 0.46 when it is not. We also observe that different teacher model
קרא במקור המקורי