כתבה
arXiv cs.LG ·
Naver-News-KO: A Korean News Summarization Dataset for Open-Source Fine-Tuning of Summarization Models
תקציר מקורי באנגליתarXiv:2607.20442v1 Announce Type: cross Abstract: We release Naver-News-KO, a Korean news summarization dataset of 27,400 (document, summary) pairs collected from Naver News over a ten-day window in July 2022 across two categories (Economy and IT/Science; 77/23 split), with train/validation/test partitions of 22,194 / 2,466 / 2,740 and a mean per-record document-to-summary character-compression ratio of 6.03x. The dataset has been publicly hosted on the Hugging Face Hub since January 2023 and, as of May 2026, receives approximately 33,000 downloads per month; community-maintained Korean summarization models fine-tuned on it include Gemma-2B-ko and Gemma2-9B variants. This technical report (i) documents the collection protocol, the column schema, and the split construction, (ii) reports cor
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית