כתבה
arXiv cs.AI ·
CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets
תקציר מקורי באנגליתarXiv:2610.07132v1 Announce Type: cross Abstract: Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית