כתבה
arXiv cs.AI ·
Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets
תקציר מקורי באנגליתarXiv:2602.14536v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have seen remarkable advancements, achieving state-of-the-art results in diverse applications. Fine-tuning, an important step for adapting LLMs to specific downstream tasks, typically involves further training on corresponding datasets. However, a fundamental discrepancy exists between current fine-tuning datasets and the token-level optimization mechanism of LLMs: most datasets are designed at the sentence-level, which introduces token-level noise, causing negative influence to final performance. In this paper, we propose XTF, an explainable token-level noise filtering framework. XTF decomposes the complex and subtle contributions of token-level data to the fine-tuning process into three distinct and ex
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית