יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning

CluSTER - פרקטיקה לקיצור זמן אימון של LLMs על ידי קבוצות ואלוקציה מאוזנת
תקציר מקורי באנגליתarXiv:2609.12584v1 Announce Type: new Abstract: Instruction-tuning datasets for large language models (LLMs) are often large, redundant, and imbalanced, limiting efficient adaptation. Naive large-batch fine-tuning repeatedly includes overrepresented sample groups while weakly covering underrepresented but informative ones, especially under data parallelism (DP) across multiple GPUs. We propose CluSTER, a Cluster-aware balanced Sampling framework for Training Efficient data Reduction in DP instruction tuning. CluSTER curates a representative reduced dataset through gradient-space clustering and DP-aware balanced allocation, ensuring dual-level coverage across clusters and workers, while preserving the original data distribution by weighted update. As a result, CluSTER reduces redundant comp
קרא במקור המקורי