כתבה
arXiv cs.CL ·
An Information-Theoretic Approach to Identifying Formulaic Clusters in Textual Data
תקציר מקורי באנגליתarXiv:2503.07303v3 Announce Type: replace Abstract: Texts, whether literary or historical, exhibit structural and stylistic patterns shaped by their purpose, authorship, and cultural context. Formulaic texts, which are characterized by repetition and constrained expression, tend to differ in their \textit{information content} (as defined by Shannon) compared to more dynamic compositions. Identifying such patterns in historical documents, particularly multi-author texts like the Hebrew Bible, provides insights into their origins, purpose, and transmission. This study aims to identify formulaic clusters: sections exhibiting systematic repetition and structural constraints, by analyzing recurring phrases, syntactic structures, and stylistic markers. However, distinguishing formulaic from non-
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית