כתבה
arXiv cs.AI ·
AstroAgentBench: Evaluating Agentic Planning on Space Mission Planning Tasks
תקציר מקורי באנגליתarXiv:2601.11354v3 Announce Type: replace Abstract: Recent LLM-for-Space systems address mission planning, scheduling, operations support, simulator control, and autonomy, but their evaluations use different task contracts, control settings, simulators, and success criteria. We introduce AstroAgentBench, a seven-family benchmark for executable space mission planning in the domains of scheduling, observation planning, constellation design, and relay support. For each case, an agent submits a planning artifact that is checked by an external verifier for schema, timing, geometry, resources, and mission value. Results report validity and normalized scores, with comparisons to task-specific solver references. Across five LLM agent systems and 35 held-out cases, the strongest systems approach or
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית