יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

NeuroActiSep: איתור הזיות מעובדות בניורונים

NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass
NeuroActiSep הוא שיטה לאיתור הזיות מעובדות במודלים של שפה. השיטה משתמשת בניורונים קדמיים כדי לזהות דפוסים של אמת ועובדות. המחקר מראה כי השיטה יכולה לזהות הזיות ברמת דיוק גבוהה.
תקציר מקורי באנגליתarXiv:2609.14448v1 Announce Type: cross Abstract: Hallucination in large language models reduces their reliability and slows adoption. Various white-box studies have used internal representations to detect patterns of truthfulness and factuality. A less-studied approach is to identify feed-forward neurons correlated with hallucination. We propose a method to rank feed-forward neurons at the final prompt token using a custom neuron selection dataset. We transfer the selected neuron identities to train hallucination classifiers on other factual question answering datasets. Our work provides empirical evidence that probes trained using the features from the selected neurons perform on par with probes trained on internal states. We also analyze the distribution of selected neurons and the effe
קרא במקור המקורי