כתבה
arXiv cs.AI ·
שיפור נאמנות OCR באמצעות ניכור מדורג
Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation
GAD-RL משפר את נאמנות OCR באמצעות ניכור מדורג. המודל Qwen3.5-2B השיג 59.92% Micro Recall על CHAOS-Bench. GAD-RL מותאם לפי ביצועי המשימה והתפלגות מקומית.
תקציר מקורי באנגליתarXiv:2609.38282v1 Announce Type: new Abstract: Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Sequence-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves. Offline analysis shows that supervision from a fixed teacher becomes progressively less favorable as the student improves, both across training checkpoints and across response groups with different task rewards. Motivated by this observation, we introduce GAD-RL, which adaptively regulates teacher supervision during joint post-training according to the student's current task performance and local distributions. A frozen teacher condi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית