יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מרכוש מודלי שפה גדולים לגרף מנבאים

Distilling LLM Reasoning into Graph of Concept Predictors
פותח כלי חדש להפחתת עלויות ושיפור ביצועים של מודלי שפה גדולים. הכלי, Graph of Concept Predictors, מאפשר למודלים ללמוד ממורים ולשפר את יכולתם לנבא מושגים. הפיתוח יכול לשפר את ביצועי המודלים תוך הפחתת עלויות ושיפור יציבות האימון.
תקציר מקורי באנגליתarXiv:2602.03006v3 Announce Type: replace Abstract: Deploying Large Language Models (LLMs) for discriminative workloads is often limited by inference latency, compute, and API costs at scale. Active distillation reduces these costs by querying an LLM oracle to train small discriminative students, but most pipelines distill only final labels, discarding intermediate reasoning signals and offering limited diagnostics of what reasoning is missing and where errors arise. We propose Graph of Concept Predictors (GCP), a reasoning-aware active distillation framework in which the teacher's reasoning is elicited as a directed acyclic graph of intermediate concepts and mirrored in the student. GCP enhances sample efficiency through a graph-aware acquisition strategy that weights per-concept uncertai
קרא במקור המקורי