יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה MarkTechPost ·

Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks

Datalab השיקה OmniExtractBench, במבחן פתוח להפרעת תיעוד רצופה.
תקציר מקורי באנגליתDatalab has released OmniExtractBench , an open benchmark for structured document extraction. It tests how accurately a system fills a JSON schema from a PDF. The benchmark pools 620 documents from 4 existing benchmarks. One deterministic scorer grades all of them and explains each decision. The release lands while extraction vendors publish their own leaderboards. Datalab argues those leaderboards are hard to compare or audit. OmniExtractBench is its attempt at a shared yardstick. Is it deployable? Yes , the scorer installs from PyPI as omni-extract-bench (v0.1.7, Python 3.11+, SciPy only) under Apache 2.0. Rerunning vendors requires your own API keys and paid credits. What is OmniExtractBench? OmniExtractBench is a structured extraction benchmark built by Datalab. Each task gives a syste
קרא במקור המקורי