כתבה
arXiv cs.AI ·
A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation
תקציר מקורי באנגליתarXiv:2609.01315v2 Announce Type: replace Abstract: Building an omni-modal foundation model means evaluating it across text, image, video, and audio. Excellent evaluation toolkits exist for each modality, but their inference engines, prompt conventions, and metric implementations are mutually incompatible, so practitioners end up maintaining separate environments for every toolchain and still struggle to compare results across them. OmniEvaluator grew out of this need in our own model development: rather than reimplementing benchmarks, it connects existing inference engines and curated evaluation libraries at a higher level, exposing four inference backends, four evaluation frameworks, and over a thousand benchmarks through a single interface. Every run is recorded as an artifact capturing
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית