יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

SHELF: תשתית סינתטית לבדיקה של תכונות LLM

SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking
מערכת חדשה לבדיקת תכונות LLM בספריות וארכיונים. SHELF מאפשרת יצירת נתוני בדיקה ושיטות בדיקה עבור LLMs. המערכת כוללת 62,899 דוקומנטים שנכתבו על ידי LLMs ומכילה 5 משימות שונות. המערכת נועדה לספק תשתית לבדיקת תכונות LLM בספריות וארכיונים.
תקציר מקורי באנגליתarXiv:2609.03047v1 Announce Type: new Abstract: Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmarks do not systematically test their bibliographic work. They need to know which methods work for their tasks and what those methods require to run. SHELF, the Synthetic Harness for Evaluating LLM Fitness, addresses this gap. It is a Python system that turns labelled taxonomies, writing specifications, and a generation budget into controlled benchmark data and evaluation tasks. This first release contains 62,899 model-written documents based on Library of Congress vocabularies, with tasks for classification, clustering, retrieval, pair classification, and instruction retrieval. We compare TF, TF-IDF, BM25, popular encoders, and, on subj
קרא במקור המקורי