כתבה
arXiv cs.CL ·
מפיקים מפיקסלים לזוגות: בניית מבחן רב-מודלית להפקת זוגות ערך-ערך בתנאי תעתוע
From Pixels to Pairs: A Comprehensive Benchmark of LLM-Driven Key-Value Extraction in Noisy Document Settings
במאמר זה, נבחן את יכולתם של רשתות לשון (LLMs) להפיק זוגות ערך-ערך מתיקוני תעתוע. נבחן 136 תצורות ניסויים ו-17,688 תחושות-דוקומנט עם פיזור דטרמיניסטי. המבחן כולל תוצאות זיהוי-שבירה, תוצאות-מקביל, ובדיקת רגישות. התוצאות מפריכות שלושה תפיסות יסוד: (1) תפיסה נקי-טקסט יכולה לחזות את רובות-עולם, (2) דירוגי המודלים נשארים קבועים בין טקסט-זהב וטקסט-OCR, ו(3) הוספת תרגילים קצרים יכולה לשפר את דיוק ההפקה. המבחן נכתב בשביל תחקיר רפודובילי.
תקציר מקורי באנגליתarXiv:2609.17538v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated strong capabilities in document key-value pair (KVP) extraction, yet controlled evaluations of their robustness to optical character recognition (OCR) output remain limited. This leaves an important gap in understanding their reliability in real-world OCR-to-LLM pipelines. Unlike end-to-end Vision-Language Models (VLMs), which jointly perform visual perception and semantic extraction, modular pipelines allow these stages and their errors to be isolated and audited. We introduce a controlled benchmark that distinguishes downstream LLM extraction behavior from upstream OCR degradation. It evaluates 136 experimental configurations and 17,688 document-level inferences generated with deterministic
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית