כתבה
arXiv cs.LG ·
DataPrep-Bench: בנצ'מרק ל-LLM
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
DataPrep-Bench הוא בנצ'מרק חדש ל-LLM. הוא בודק יכולת הכנת נתונים ואיכותם. המחקר משווה שיטות שונות, כולל Data-Construction-Skill, ומציג את DAS, מדד לאיכות נתונים.
תקציר מקורי באנגליתarXiv:2607.20465v1 Announce Type: new Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end. We view LLM-driven data preparation as comprising two complementary capabilities: data construction, which transforms raw sources into supervised training data, and data quality evaluation, which predicts the training value of candidate datasets before downstream training; throughout, "quality" refers to downstream training utility rather than surface-level textual properties. We introduce DataPrep-Bench, the first unified benchmark that jointly evaluates both capabilities under a shared downstream-grounded
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית