כתבה
arXiv cs.CL ·
Useful Features, Backward Scores: OOD in Language-Model Trajectories
תקציר מקורי באנגליתarXiv:2605.00269v2 Announce Type: replace Abstract: Out-of-distribution (OOD) detectors prioritize inputs for closer inspection. Yet features that distinguish input groups need not yield a useful anomaly ranking. We analyze this gap in language-model trajectories under text-length control and fixed score directions. On Spam development data, an input adaptation of D^2HScore falls from raw AUROC 0.919 to 0.530 after length matching. On length-matched, held-out HateSpeech inputs, the same features yield AUROC 0.644 for a labeled linear classifier but 0.444 for an ID-fitted distance score. ToxicChat shows the same contrast. Feature-selection and backbone controls retain the main reversal pattern. Frozen Civil Comments and TweetEval irony tests also reverse (0.467 and 0.435), extending the fin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית