יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

CausalVerify: בדיקת עמידה לבדיקה לעבודות חישוב סיבתיות של LLM

CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows
CausalVerify מבצע בדיקות עמידה לבדיקה לעבודות חישוב סיבתיות של LLM. הבדיקות כוללות 259 מאמרי כלכלה ו-100 סצנריות סינתטיות.
תקציר מקורי באנגליתarXiv:2609.07944v1 Announce Type: new Abstract: Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate. CausalVerify studies this verification problem for structured econometric causal-estimation workflows by separating realistic interpretation from verifiable computation. It pairs 259 published economics papers (reconstructed research question, data description, institutional context) with 100 fixed-seed synthetic scenarios that realise CSV datasets for difference-in-differences, event study, instrumental variables, and regression discontinuity designs. Experiment A (real-paper text agreement) scores method-family and direction agreement against four-LLM consensus
קרא במקור המקורי