כתבה
arXiv cs.CL ·
MetroLLM-Bench: בדיקת מודלים שפה כרכיב תחבורה
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
MetroLLM-Bench הוא בנץ'מרק לבדיקת מודלים שפה כרכיב תחבורה. הוא כולל 955 מקרים שונים, ובודק מודלים כגון Qwen ו-GPT. הבדיקה כוללת נושאים כגון תכנון נסיעות, חישוב תעריפים ונגישות.
תקציר מקורי באנגליתarXiv:2609.10016v1 Announce Type: cross Abstract: We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-quality components form Tier 2, six of which use a language-model judge. We report Tier 1 and the combined score of both tiers. A stratified 75/25 split reserves 717 cases for training-data generation
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית