כתבה
arXiv cs.LG ·
בדיקת תקן: דגימות סיסמאות מודלים נגד מודלי מחלקות מאומנים ומודלי שפה לשערי החלטה אוטומטיים
Benchmarking System One decision models against trained classifiers and language models for automated decision gates
בדיקת תקן של מודלי סיסמאות מודלים נגד מודלי מחלקות מאומנים ומודלי שפה לשערי החלטה אוטומטיים. המחקר משווה את תפוקתם של מודלי סיסמאות מודלים, מודלי מחלקות מאומנים ומודלי שפה בשערי החלטה אוטומטיים.
תקציר מקורי באנגליתarXiv:2610.00346v2 Announce Type: replace Abstract: Software that hands branching decisions to a model needs a declared option and a probability it can threshold. Typed decision models, also called System One models, return such probabilities without generating text, while supervised classifiers and generative language models are the established alternatives. One harness sends eight decision-model checkpoints from six families, including the hosted model Jev, and four open generative models from three developers the same semantic requests, and scores trained and zero-shot classifiers on the same workflow, intent, emotion and social-science items. With task labels, a fine-tuned DeBERTa-v3-large has the highest observed accuracy on every labeled benchmark but one. Without labels, no decision
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית