כתבה
arXiv cs.LG ·
FastKernels: Benchmarking GPU Kernel Generation in Production
תקציר מקורי באנגליתarXiv:2605.23215v2 Announce Type: replace Abstract: LLM-based agents for GPU kernel generation are advancing rapidly, but the benchmarks they optimize against evaluate kernels in isolation, with synthetic inputs and weak baselines, rewarding sandbox speedups that break or vanish in real inference systems. We introduce FastKernels, a benchmark of 384 tasks drawn from 47 representative architectures across 8 categories, whose kernels suffice to reimplement 94.6% (472/499) of HuggingFace Transformers architectures with outputs matching the native implementations. Each task mirrors the interface of the corresponding production module and is scored against the kernels production frameworks ship, and tasks form a compositional hierarchy, from primitives to full models, in which higher-level modu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית