כתבה
arXiv cs.AI ·
Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems
תקציר מקורי באנגליתarXiv:2609.05928v1 Announce Type: cross Abstract: Large language models now compute correct tax liabilities on over 90% of well-formed cases in statutory benchmarks, which makes them candidates for the tax-advisory and compliance systems that consume such an answer directly. Real legal inputs, however, are frequently defective: required facts are missing, or stated facts contradict one another. Accuracy on clean benchmarks says nothing about how a model behaves then, and a system that computes straight through a defective input returns a confident number with no sign that anything is wrong. This raises two questions: does a model asked to solve a case abstain when the input is defective, and when it does not, can the same model catch the defect when asked instead to verify the input? We st
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית