כתבה
arXiv cs.CL ·
Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable
תקציר מקורי באנגליתarXiv:2607.21340v2 Announce Type: replace Abstract: In capital-markets workflows the question is rarely whether a large language model can produce a fluent draft, but whether the draft is bankable: defensible in front of a counter-party or a regulator, with the documents in hand. Existing methods address parts of that gap: open-domain QA benchmarks reward surface accuracy, and finance benchmarks (FinanceBench, FinQA, ConvFinQA) advance document-grounded and numerical QA but evaluate at the question-answer layer rather than the workflow outputs practitioners defend. We introduce CM-LRS, a Capital Markets LLM Reliability Score, evaluating outputs at the workflow-output layer across seven dimensions: factual accuracy, evidence traceability, numerical consistency, workflow completeness, source
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית