יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

DI-Bench: יצירת בנצ'מרקים למודלי נתונים לצורך חקירת נתונים בעסקים

DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents
DI-Bench מפיקה בנצ'מרקים ממומשים למשימות חקירת נתונים. הבנצ'מרקים כוללים שאלות שמערבות חישובים וחיפוש ידע. נבחנו ארבעה מודלים, והתגלה כי הם הגיעו לדיוק של 32% במשימות חישוביות שבהן כללי עסקים משנים את החישוב.
תקציר מקורי באנגליתarXiv:2609.05776v1 Announce Type: new Abstract: Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical computation, and constructing such benchmarks manually is costly. We present DI-Bench, a pipeline for generating realistic benchmarks for data intelligence (DI), the practice of extracting insights from large volumes of enterprise data. To emulate realistic DI tasks that require both computation and knowledge retrieval, DI-Bench builds an artifact linkage graph over data tables, dimensions, metrics, and documents to form questions involving structured data and associated knowledge. Ground truth answers are derived via query execution, followed by LLM question generation a
קרא במקור המקורי