יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אחד גבול אינו תואם לכל שפות: תקיפה-שפה-תנאית לסווג טקסט נמוך-משאבים באופן אמין וקצר-זמן

One Threshold Does Not Fit All Languages: Language-Conditional Deferral for Reliable and Efficient Low-Resource Text Classification
במאמר זה, המחברים מציעים פתרון לסווג טקסט נמוך-משאבים באופן אמין וקצר-זמן. הם מציעים להשתמש בתקיפה-שפה-תנאית, שהיא טכניקה שמאפשרת להעריך את הביטחון של המודל באופן שונה לכל שפה. המחברים מדגימים את הפתרון באמצעות ניסויים שנערכו ב-16 שפות אפריקאיות ו-12 שפות נוספות.
תקציר מקורי באנגליתarXiv:2609.37861v1 Announce Type: new Abstract: In the Global South, the lower-income countries of Africa, Asia, and Latin America where most of the world's languages are spoken, a deployed text classifier usually runs on ordinary CPUs, serves many languages with a single model, has few labeled examples in any of them, and relies on people to catch its mistakes. Such a system is only useful if it can promise how often it will be wrong: at most a fixed fraction of the labels it assigns on its own may be incorrect, and everything else must go to a person. Split conformal prediction delivers this promise through a single confidence threshold, normally estimated on validation data pooled across languages. We ask whether the promise reaches every language, and it does not. On MasakhaNEWS (16 Af
קרא במקור המקורי