כתבה
arXiv cs.CL ·
LentEx: כריית ישויות לטנטיות
LentEx: Generalizable Latent Entity Extraction via Synthetic Data and Instruction-Tuned LLMs
LentEx היא שיטה חדשה לכריית ישויות לטנטיות, המשתמשת בנתונים סינתטיים ובהדרכה מותאמת למודלים גדולים של שפה. השיטה משפרת באופן משמעותי את הביצועים במגוון משימות, ומוכיחה עצמה כמוצלחת ביישומים מעשיים.
תקציר מקורי באנגליתarXiv:2609.04511v1 Announce Type: new Abstract: Latent entity extraction (LEE) tackles the challenge of identifying implicit, contextually inferred entities within free text-an area where traditional entity extraction methods fall short. In this paper, we introduce LentEx, a novel framework for latent entity extraction that leverages synthetic data generation and instruction fine-tuning to optimize smaller, efficient large language models (LLMs). Latent entities, which are often abstract and thematic, are crucial for applications such as retrieval-augmented generation (RAG), customer persona analysis, and knowledge graph enrichment. LentEx addresses the scarcity of labeled datasets by employing a template-based approach to generate diverse, contextually rich synthetic data, ensuring high v
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית