כתבה
arXiv cs.AI ·
dattri-LLM: ספריית יחידה ומציאותית לאטריבוציה של נתוני הכשרה בסקאלה של LLM
dattri-LLM: A Unified and Efficient Library for Training Data Attribution at LLM Scale
dattri-LLM היא ספריית יחידה ומציאותית שמאפשרת אטריבוציה של נתוני הכשרה בסקאלה של LLM. היא משתמשת בייצוגים קצרים של גרדיאנטים ומסלולים דינמיים לצורך רוטציה של פעולות גרדיאנט. הספרייה תומכת במגוון של שיטות אטריבוציה, כולל שיטות דומיננטיות, שיטות השפעה-מעקב, ושיטות מסלוליות. dattri-LLM תומכת גם בשיטות של בחירת נתונים בזמן הכשרה.
תקציר מקורי באנגליתarXiv:2609.38767v1 Announce Type: cross Abstract: Training data attribution (TDA) estimates the contribution of individual training examples to model outputs. Most scalable TDA methods rely on per-example gradients, whose computation and use at LLM scale pose challenges in efficiency, compatibility, and extensibility. We introduce dattri-LLM, a TDA library that makes gradient-based attribution more practical at scale. For efficiency, dattri-LLM uses compact gradient representations and dynamically routes gradient operations based on a cost model. For compatibility, its capture mechanism collects per-example gradients from existing training loops that call backward(), without requiring changes to the loop or its configuration. This includes distributed training with DDP and FSDP and pipelin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית