יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ScienceClaw: תקיפת עצמית רציפה של סוכני AI למדע

ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
ScienceClaw היא תקיפת עצמית רציפה של סוכני AI למדע, המבוקרת ב-23 תחומים. הפרויקט נועד לבחון את יכולתם של סוכני AI לשפר את עצמם באופן רציף, ולבדוק את יכולתם להתאים לשינויים בתחומי המדע. קוד הפרויקט זמין בגיטהב.
תקציר מקורי באנגליתarXiv:2610.08691v1 Announce Type: new Abstract: Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates. ScienceClaw-Eval spans 23 disciplines and measures scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost through sequential streams and independent reset evaluation. Our framework repairs executable workflows through multi-turn interaction, converts re-execution-verified failure--success trajectories into
קרא במקור המקורי