כתבה
arXiv cs.CL ·
PROMPT2BOX: שיפור גילוי חולשות LLM
PROMPT2BOX:Improving LLM Weakness Discovery and Specificity Estimation by Uncovering Entailment Structure among Prompts
PROMPT2BOX הוא כלי לגילוי חולשות במודלים LLM. הוא משתמש באימון מיוחד כדי ליצור מרחבי עיגול לקידומי טקסט, מה שמאפשר ניתוח רגיש יותר של חולשות המודל. הכלי הראה יכולת לזהות 13.5% יותר חולשות LLM מאשר שיטות קודמות.
תקציר מקורי באנגליתarXiv:2603.21438v3 Announce Type: replace Abstract: To discover the weaknesses of LLMs, researchers often embed prompts into a vector space and cluster them to extract insightful patterns. However, vector embeddings primarily capture topical similarity; as a result, prompts that share a topic but differ in specificity, and consequently in difficulty, are often represented similarly, making fine-grained weakness analysis difficult. To address this limitation, we propose Prompt2Box, which embeds prompts into a box embedding space using a trained encoder. The encoder, trained on existing and synthesized datasets, outputs box embeddings that capture not only semantic similarity but also specificity relations between prompts (e.g., "writing an adventure story" is more specific than "writing a s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית