יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אותות הגברה לניתוב LLM

Evaluating Escalation Signals for LLM Routing: Targets, Controls, and Five Ways to Fool Yourself
חוקרים בדקו אותות הגברה לניתוב LLM, כולל אנטרופיה סמנטית. הם מצאו שאות זה יעיל ב-GSM8K, אך לא בבנך'ים אחרים. הם גם הציעו רשימת בדיקות למניעת תוצאות שווא.
תקציר מקורי באנגליתarXiv:2610.07354v1 Announce Type: new Abstract: Deciding when to escalate a query from a small language model to a larger one requires a cheap signal that predicts, before the large model is called, whether escalating would help. Semantic entropy, originally developed to detect hallucinations, is a natural candidate: it measures how much a model's sampled answers disagree in meaning, and high disagreement often signals an unreliable answer. We test it across three benchmarks and two model families. On GSM8K, with a small/large pair about twelve times apart in size, semantic entropy reliably distinguishes the small model's mistakes (AUROC 0.871) and improves routed accuracy over random escalation by up to nine points at matched cost. An earlier strong-looking result on a synthetic benchmark
קרא במקור המקורי