כתבה
arXiv cs.AI ·
DepthBenchCAD: כאשר ניתוח עמוק יותר יוצר תוצאות יותר נאמנות?
DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?
ניתוח עמוק יותר של תוכנות עשוי לא להביא לתוצאות יותר נאמנות. המאמר עוסק בבחינה של תוכנות CAD ובאיך ניתוח עמוק יותר יוצר תוצאות יותר נאמנות. התוצאות עולות בקנה מידה ומראות שהעומק של הניתוח יכול להיות תלוי במקור של החוסר בוודאות.
תקציר מקורי באנגליתarXiv:2609.15122v1 Announce Type: cross Abstract: Generative CAD models are expected to remain behaviorally correct after parameter edits, so increasing the number of edit checks is often treated as a direct route to more reliable evaluation. Under a fixed budget, however, auditing each program more thoroughly reduces the number of tasks and independent generations that can be evaluated, which can ultimately make model-level estimates less accurate. We study this phenomenon and the conditions under which it arises. We decompose behavioral evaluation into three evidence levels: task templates, stochastic generations, and within-program edits. We define an average failure risk that is invariant to audit depth, and combine three-level variance with measured execution costs to analyze the trad
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית