כתבה
arXiv cs.AI ·
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
תקציר מקורי באנגליתarXiv:2605.04135v4 Announce Type: replace-cross Abstract: LLM evaluations in applied domains tend to reflect models that were already outclassed at time of publication. We observe a publication elicitation gap: the distance between the AI systems generating the results reported in an academic paper and the AI systems that a current reader of that paper would reasonably assume are being referenced. We systematically sweep OpenAlex from 2022-01-01 to 2026-04-01 (n = 112,303 LLM keyword matches). Then, we identify what models were evaluated (n = 18,574 admissible records). We then rank each evaluated LLM against a frontier LLM based on the Epoch AI Capabilities Index (ECI), an aggregate LLM capability score. We find that the median paper's models are worse than the frontier LLM at the time of
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית