כתבה
arXiv cs.AI ·
HARISSA: בדיקות-עצמן בזמן השקעה להקטנת זמן ושיפור בטיחות של פריסת מודלי שפה
HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment
HARISSA מציעה בדיקות-עצמן בזמן השקעה להקטנת זמן ושיפור בטיחות של פריסת מודלי שפה. השיטה, המכונה HARISSA, עושה שימוש במצבי הסתר של המודל כדי לחזות האם המודל יענה כראוי ואם התשובה שהוא נותן היא נכונה.
תקציר מקורי באנגליתarXiv:2609.38006v1 Announce Type: new Abstract: Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stays local faces two decisions for hard queries instead. First, it can spend more computation on a query, e.g., reasoning before answering, which raises accuracy at a cost in latency, so it must decide which queries are worth the extra computation (efficiency). Second, some queries are beyond the local model, and delivering a wrong answer is worse than deferring the query to a human in the loop, so it must decide which answers are
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית