כתבה
arXiv cs.AI ·
גישור פער המודליות באודיו קליני ארוך: מחקר השוואתי
Bridging the Modality Gap in Long-Form Clinical Audio: A Comparative Study of Lightweight and Heavyweight End-to-End SOAP Generation
צוות ASLP מציג מערכת רב-מודלית ליצירת תיעוד רפואי מאודיו. המערכת משתמשת בגישה E2E ומתעלמת מתמלילים ביניים. הם בנו מאגר נתונים גדול והשתמשו באימון רב-שלבי.
תקציר מקורי באנגליתarXiv:2609.14467v1 Announce Type: cross Abstract: Automating clinical documentation from long-form doctor-patient conversations remains challenging for modern audio-language models. While cascaded ASR systems perform well, end-to-end (E2E) models often struggle with information loss and hallucinations on extended audio. For the BeTraC 2026 challenge, the ASLP team presents a fully E2E multimodal system that generates structured SOAP notes directly from audio, bypassing intermediate transcripts. We constructed a 1.41-million-sample multi-task corpus and applied a multi-stage pipeline: domain pre-training, supervised fine-tuning, and reward optimization. Evaluating the architecture under both Lightweight (3B) and Heavyweight (30B) constraints reveals that each training stage progressively en
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית