כתבה
arXiv cs.LG ·
ArrivalBench: פייפלינס של נגיני גנרטיביים נכונים פעם אחת ולא נכונים תחת זמן
ArrivalBench: Agent-Generated Data Pipelines Are Correct Once and Wrong Under Time
במבחנים של פייפלינס שנוצרו על ידי נגיני, נמצאו פערים בין תוצאות שנקבעו פעם אחת לבין תוצאות שנקבעו תחת תקופות שונות
תקציר מקורי באנגליתarXiv:2610.02363v1 Announce Type: new Abstract: Benchmarks for agent-generated data work grade a pipeline by running it once against a fixed snapshot. ArrivalBench instead re-executes the pipeline an agent leaves behind under adversarial but replayable delivery schedules (late, duplicated, out-of-order and retried records) and requires its final state to equal a batch recomputation of the complete log. Because the oracle recomputes rather than classifies, a wrong table and a crash are distinct verdicts: a crash is visible to monitoring a team already runs, and a wrong table is not. On 40 tasks we built, our reimplementation of single-execution grading certifies 86-100% of the pipelines eleven models produce; re-executing the same artifacts finds 7.0-79.2% of the certified ones silently wro
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית