יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Useful Features, Backward Scores: OOD in Language-Model Trajectories

תקציר מקורי באנגליתarXiv:2605.00269v2 Announce Type: replace Abstract: Out-of-distribution (OOD) detectors prioritize inputs for closer inspection. Yet features that distinguish input groups need not yield a useful anomaly ranking. We analyze this gap in language-model trajectories under text-length control and fixed score directions. On Spam development data, an input adaptation of D^2HScore falls from raw AUROC 0.919 to 0.530 after length matching. On length-matched, held-out HateSpeech inputs, the same features yield AUROC 0.644 for a labeled linear classifier but 0.444 for an ID-fitted distance score. ToxicChat shows the same contrast. Feature-selection and backbone controls retain the main reversal pattern. Frozen Civil Comments and TweetEval irony tests also reverse (0.467 and 0.435), extending the fin
קרא במקור המקורי