יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

MetroLLM-Bench: בדיקת מודלי שפה כמערכת ריצה של קיוסקי תחבורה

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
בדיקת מודלי שפה כמערכת ריצה של קיוסקי תחבורה. המאמר מציג תקן לבדיקת מודלי שפה כמערכת ריצה של קיוסקי תחבורה, ומדגים את יעילותם של מודלי שפה כקיוסקי תחבורה.
תקציר מקורי באנגליתarXiv:2609.10016v1 Announce Type: cross Abstract: We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-quality components form Tier 2, six of which use a language-model judge. We report Tier 1 and the combined score of both tiers. A stratified 75/25 split reserves 717 cases for training-data generation
קרא במקור המקורי