כתבה
arXiv cs.AI ·
MINCE: Shrinking LLM Evaluation Datasets via Few-Model Monte Carlo Calibration
תקציר מקורי באנגליתarXiv:2606.22826v2 Announce Type: replace Abstract: Evaluating LLMs across many model variants---quantized, fine-tuned, or deployment-specific---requires running large benchmarks repeatedly, a process that can take tens of hours per model on edge hardware such as NPUs. Existing subset selection methods reduce this cost but depend on large calibration pools or learned prediction layers. We introduce MINCE (Monte Carlo Informed N-sizing for Compact Evaluation), which uses Monte Carlo simulation over per-item logs from a small set of calibration models to find the minimum subset size that bounds accuracy drift and then fixes a randomly sampled subset at that size, with no prediction layer needed. MINCE reduces IFEVAL by 54\%, MMLU by 89\%, GSM8K by 70\%, and MMLU-Pro by 88\% with maximum drif
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית