כתבה
arXiv cs.LG ·
Utility-Aware Data Pricing: Token-Level Quality and Empirical Training Gain for LLMs
תקציר מקורי באנגליתarXiv:2604.22893v2 Announce Type: replace Abstract: Traditional ``row-count $\times$ quality coefficient'' approaches fail to capture the nonlinear utility of data for Large Language Model (LLM) capabilities. We present a utility-aware data valuation framework that moves from static accounting toward utility-based pricing. The framework combines token-level information density and data quality, empirical training contribution estimated by influence functions, proxy models, and Data Shapley values, and cryptographic mechanisms based on hash commitments, Merkle trees, and tamper-evident training ledgers. Preliminary evaluation includes a smoke-scale multi-domain study and a synthetic duplication test. Results show that proxy-based empirical gain achieves strong ranking agreement with proxy-d
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית