כתבה
arXiv cs.AI ·
מערכת טקסונומית סמויה: פרקטיקה לבניית מודלי זיהוי ובחינת החלטותיהם
The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection
מערכת טקסונומית סמויה לבניית מודלי זיהוי ובחינת החלטותיהם. המאמר מציע פרקטיקה לבניית מודלי זיהוי ובחינת החלטותיהם, כולל גם ניתוח של פרקטיקה זו. המאמר נכתב על ידי [Claude, Gemini, LangGraph, GPT-5].
תקציר מקורי באנגליתarXiv:2608.26423v2 Announce Type: replace-cross Abstract: This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be trusted. This framework, the Latent Diagnostic Taxonomy, consists of (i) constructing a dimensionality-optimized classifier, in which the embedding dimensionality is empirically selected via cross-validated performance rather than fixed a priori, (ii) locating a relatively small set of latent support vectors (~ 29% of total training examples) representing influential prompts for identifying tokens that alter the classifier's predicted labels, and (iii) utilizing such tokens and their associated attack magnitudes for constructing a diagno
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית