כתבה
arXiv cs.LG ·
TwinICL: Diagnosing Multimodal In-Context Learning through Paired Counterfactuals
תקציר מקורי באנגליתarXiv:2609.15028v1 Announce Type: cross Abstract: In-context learning (ICL) enables models to infer tasks from demonstrations, but existing benchmarks generally lack matched text and image versions needed to compare ICL performance across modalities. We introduce TwinICL, a procedurally generated benchmark providing such pairs for controlled comparison. Across six open-weight models and 38 tasks, multimodal ICL consistently underperforms text-only ICL, with gaps varying by task family. To test whether this gap can be recovered, we target visual access, task framing, and reasoning through three interventions. Their combination recovers strong multimodal ICL performance on a diagnostic subset, despite limited or inconsistent individual effects. To distinguish difficulties in executing tasks
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית