כתבה
arXiv cs.CL ·
Measuring the Creativity of Frontier LLMs in Automated Research
תקציר מקורי באנגליתarXiv:2609.14057v3 Announce Type: replace Abstract: Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated. We propose a set of metrics to evaluate creativity along the two dimensions of valueness and novelty. Valueness assesses whether each proposed idea is useful, while novelty is evaluated from three perspectives: whether the same idea has appeared before (Exact-Match P-Novelty), whether the modified variable or variable combination has been explored before (Variable-level P-Novelty), which reflects the breadth of research-space exploration, and whether the proposed idea is explicitly attributed to external knowledge in the model's reasoning (H-Novelty). Our evaluation shows that the models achi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית