כתבה
arXiv cs.AI ·
FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution
תקציר מקורי באנגליתarXiv:2609.36651v1 Announce Type: cross Abstract: Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית