כתבה
arXiv cs.AI ·
ArchitectureIQ: On the Measure of Training Intuition
תקציר מקורי באנגליתarXiv:2609.39714v1 Announce Type: new Abstract: Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuition of LLMs and humans, we introduce the ArchitectureIQ benchmark. Each question presents a synthetic dataset and several training recipes, and the test-taker is asked to predict the recipe yielding the best test metric. Overall, we find that LLMs' model intuition is good but has four limitations: (1) The intuition is imperfect, or even sub-human in some cases. Frontier models achieve around 76% accuracy (random choice 33%) vs best human researcher (66.0%), yet remain far from perfect. For architecture-only questions, best human achieves 65% while GPT-6 Astra only has 38%. (2) The intuition is e
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית