כתבה
arXiv cs.CL ·
Text-Centric Post-Training for Omni-Modal Reasoning
תקציר מקורי באנגליתarXiv:2610.02819v1 Announce Type: new Abstract: Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics reveal multi-hop reasoning difficulties despite correct answers to all corresponding single-hop questions and suggest partial decoupling in the local optimization of perception and reasoning objectives. This motivates post-training with different emphases on these capabilities. Text-only reasoning training yields gains across data sources, model scales, and families. With the best-performing text-only configuration, supervised fine-tuning followed by reinforcement learning (RL) raises Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% over the base model, outperforming the com
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית