כתבה
arXiv cs.AI ·
מחירי נתונים המודעים לתפקוד: צפיפות תגים-על-תג ורווח טיפולני מבוסס על ניסויים למודלי שפה גדולים
Utility-Aware Data Pricing: Token-Level Quality and Empirical Training Gain for LLMs
מחירי נתונים המודעים לתפקוד למודלי שפה גדולים. המאמר מציג פרקטיקה להערכת ערך נתונים שמשלבת צפיפות תגים ואיכות נתונים, ומדדי רווח טיפולני. התוצאות מציגות רמת תאימות גבוהה עם הערכת ערך נתונים מבוססת על פרוקסי.
תקציר מקורי באנגליתarXiv:2604.22893v2 Announce Type: replace-cross Abstract: Traditional ``row-count $\times$ quality coefficient'' approaches fail to capture the nonlinear utility of data for Large Language Model (LLM) capabilities. We present a utility-aware data valuation framework that moves from static accounting toward utility-based pricing. The framework combines token-level information density and data quality, empirical training contribution estimated by influence functions, proxy models, and Data Shapley values, and cryptographic mechanisms based on hash commitments, Merkle trees, and tamper-evident training ledgers. Preliminary evaluation includes a smoke-scale multi-domain study and a synthetic duplication test. Results show that proxy-based empirical gain achieves strong ranking agreement with p
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית