יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בדיקה של נפחדות סינתטית בדגמי שפה גדולים: נפחדות סינתטית בדגמי שפה גדולים

Linear Separability of Activation Representations after Supervised Fine-Tuning on Incorrect Responses: A Study of Synthetic Dishonesty in Large Language Models
במאמר זה, חוקרים בדקו את השפעת תרגול סופרוויזד על דגמי שפה גדולים שנלמדו להפיק תשובות שקריות. הם גילו שהתרגול יוצר תבנית רצפית בפעולות הפעולה של הדגמים, שניתן לזהות. המחקר חשוב להבנת כיצד דגמי שפה גדולים עובדים וכיצד ניתן להגן עליהם מפני תשובות שקריות.
תקציר מקורי באנגליתarXiv:2605.30381v2 Announce Type: replace Abstract: When a language model is fine-tuned to produce systematically incorrect responses, does this training leave a structured, linearly recoverable trace in its internal activations? We study this question in a controlled model-organism setting using five transformer architectures spanning 1.4 to 9 billion parameters. For each model, an "honest" and a "dishonest" LoRA fine-tuned variant are constructed using identical question distributions but correct versus plausible-but-incorrect answers. Linear probes reach near-ceiling separability (AUC >= 0.9997) within the first few layers in four of five architectures. This separability transfers from TruthfulQA to held-out MMLU subjects for four models, but substantially less so for Pythia-1.4B. Six g
קרא במקור המקורי