כתבה
arXiv cs.CL ·
בדיקת היצירתיות של LLMs חדשים במחקר אוטומטי
Measuring the Creativity of Frontier LLMs in Automated Research
במאמר זה, נבדיק את היצירתיות של LLMs חדשים במחקר אוטומטי. המחקר כולל בדיקה של חמשת ה-LLMs החדשים, כולל LangGraph, Gemini ו-GPT-5.
תקציר מקורי באנגליתarXiv:2609.14057v1 Announce Type: new Abstract: Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated. In this paper, we propose a set of metrics to evaluate creativity along the two dimensions of valueness and novelty. Valueness assesses whether each proposed idea is useful, while novelty is evaluated from three perspectives: whether the same idea has appeared before (Exact-Match P-Novelty), whether a previously unexplored variable or variable combination is explored (Variable-level P-Novelty), and whether the idea directly follows retrieved external knowledge or departs from it (H-Novelty). Our evaluation shows that the models achieve relatively similar scores on most creativity metrics, but dif
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית