יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

חידוש בגילוי מונחים ללא סיווג: השבתת תפוצה זיפפיאנית

Recovering the Zipfian Distribution in Unsupervised Term Discovery
במאמר זה, המחברים מציגים חידוש בגילוי מונחים ללא סיווג, על ידי שימוש בקבוצת גרפים. הם מראים כי קבוצת גרפים משיגה תפוצה זיפפיאנית טובה יותר מאשר שיטות קבוצתיות מרכזיות.
תקציר מקורי באנגליתarXiv:2606.10781v2 Announce Type: replace-cross Abstract: Unsupervised term discovery involves segmenting unlabelled speech into word- or syllable-like units and clustering these into a lexicon of candidate types. True lexicons follow a Zipfian distribution, yet the dominant centre-based clustering approach -- K-means -- produces a more uniform distribution due to an inductive bias toward spherical clusters. In this paper we revisit graph-based clustering as a bottom-up alternative, where segment embeddings are connected by pairwise similarity and partitioned using the Leiden algorithm. We show that graph clustering substantially outperforms centre-based approaches (K-means, GMM, BIRCH) in both word- and syllable-level lexicon discovery across three languages, producing more Zipf-like dist
קרא במקור המקורי