כתבה
arXiv cs.AI ·
Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering
תקציר מקורי באנגליתarXiv:2505.13520v2 Announce Type: replace-cross Abstract: Textbook question answering (TQA) is a complex task, requiring the interpretation of complex multimodal context. Although recent advances have improved overall performance, they often encounter difficulties in educational settings where accurate semantic alignment and task-specific document retrieval are essential. In this paper, we propose a novel approach to multimodal textbook question answering by introducing a mechanism for enhancing semantic representations through multi-objective joint training. Our model, Joint Embedding Training With Ranking Supervision for Textbook Question Answering (JETRTQA), is a multimodal learning framework built on a retriever--generator architecture that uses a retrieval-augmented generation setup,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית