כתבה
arXiv cs.LG ·
Beyond Scale and Generation: Understanding Language Model-based Entity Matching
תקציר מקורי באנגליתarXiv:2607.24688v1 Announce Type: cross Abstract: Entity matching identifies records that refer to the same real-world entity. Language models can be adapted to this task through bi-encoder, cross-encoder, and generative matcher architectures. However, prior studies often conflate matcher architecture with differences in model backbone, model variant(reflecting different pretraining objectives), and model size, making it difficult to isolate the sources of performance gains. We address this issue through a controlled factorial study spanning three matcher architectures, three model variants and three model sizes from the Qwen3 family, and nine datasets, totaling 1,215 fine-tuning runs. We also evaluate cross-dataset transferability and computational cost. Our results show that model varian
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית