כתבה
arXiv cs.LG ·
BAT-CLIP: Trimodal Alignment of Brain, Audio and Text
תקציר מקורי באנגליתarXiv:2609.31180v1 Announce Type: cross Abstract: Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently multimodal speech processing. This induces a trade-off: audio anchoring preserves temporal structure but weakens linguistic separability, while text anchoring captures semantics yet discards acoustic detail. We propose BAT-CLIP, the first CLIP-style trimodal alignment framework for iEEG that jointly aligns neural embeddings to both pretrained audio and text anchors in a shared, frozen audio-text manifold. On the naturalistic Podcast benchmark, BAT-
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית