כתבה
arXiv cs.AI ·
D2K-Bench: האם סוכני LLM יכולים להפוך תכנונים מומחים לקורנלים GPU יעילים?
D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?
סוכני LLM נתקלים בקושי לייצר קורנלים GPU יעילים בלי הנחיות מומחים.
תקציר מקורי באנגליתarXiv:2610.03226v1 Announce Type: cross Abstract: GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents translate expert design guidance into efficient GPU kernels. The guidance covers L1: high-level algorithmic insights, L2: dataflow design, and L3: low-level optimization tricks, including dependencies among these levels. Pairwise runs with and without guidance share task descriptions, workloads, tools, hardware, and a 350-turn budget. Complementary assessments examine independently proposed designs and the design properties i
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית