כתבה
arXiv cs.CL ·
Zipbench: פלטפורמה זולה להקטנת מבחני בקרה רחבי-היקף של מודלי שפה גדולים
Zipbench: Low-Cost Framework for Compressing Comprehensive Benchmarks of Large Language Models
פלטפורמה חדשה להקטנת מבחני בקרה של מודלי שפה גדולים. זולה, פשוטה ומבטיחה תוצאות. כוללת גם קובץ קצר של 100+ מבחני בקרה.
תקציר מקורי באנגליתarXiv:2609.12475v1 Announce Type: new Abstract: Comprehensive benchmark suites are essential for improving large language models (LLMs), but many widely used benchmarks are redundant, making evaluation unnecessarily expensive. Although recent benchmark compression methods (BCMs) can mitigate this cost, many strong BCMs rely on large collections of per-sample evaluation results from numerous LLMs to identify representative samples. Building such collections is also expensive unless they are already public, making these methods difficult to extend to newly released benchmarks. To address this challenge, we present ZipBench, a simple and low-cost BCM with theoretical error and rank-consistency guarantees. ZipBench evaluates only a small set of anchor LLMs, synthesizes pseudo evaluation result
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית