כתבה
arXiv cs.CL ·
אחד גבול אינו תואם לכל שפות: תקיפה-שפה-תנאית לסווג טקסט נמוך-משאבים באופן אמין וקצר-זמן
One Threshold Does Not Fit All Languages: Language-Conditional Deferral for Reliable and Efficient Low-Resource Text Classification
במאמר זה, המחברים מציעים פתרון לסווג טקסט נמוך-משאבים באופן אמין וקצר-זמן. הם מציעים להשתמש בתקיפה-שפה-תנאית, שהיא טכניקה שמאפשרת להעריך את הביטחון של המודל באופן שונה לכל שפה. המחברים מדגימים את הפתרון באמצעות ניסויים שנערכו ב-16 שפות אפריקאיות ו-12 שפות נוספות.
תקציר מקורי באנגליתarXiv:2609.37861v1 Announce Type: new Abstract: In the Global South, the lower-income countries of Africa, Asia, and Latin America where most of the world's languages are spoken, a deployed text classifier usually runs on ordinary CPUs, serves many languages with a single model, has few labeled examples in any of them, and relies on people to catch its mistakes. Such a system is only useful if it can promise how often it will be wrong: at most a fixed fraction of the labels it assigns on its own may be incorrect, and everything else must go to a person. Split conformal prediction delivers this promise through a single confidence threshold, normally estimated on validation data pooled across languages. We ask whether the promise reaches every language, and it does not. On MasakhaNEWS (16 Af
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית