כתבה
arXiv cs.AI ·
בנצ'מרק רדאר: מסד נתונים ומנוע חיפוש לבחינות AI
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
בנצ'מרק רדאר הוא מסד נתונים ומנוע חיפוש לבחינות AI, המכסה מודלי שפה גדולים ובחינות אחרות. המערכת מאפשרת חיפוש וגילוי של בחינות AI, כולל מודלים כמו LLaMA.
תקציר מקורי באנגליתarXiv:2609.11115v1 Announce Type: new Abstract: Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The system combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog, mentions in model cards and technical reports, and score histories. It retains source identities and citations so readers can inspect candidate benchmarks and their evaluatio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית