כתבה
arXiv cs.AI ·
Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift
תקציר מקורי באנגליתarXiv:2607.17384v2 Announce Type: replace Abstract: This paper provides an experimentally verified formal law for calculating the uplift that diversity of thought provides in Large Language Model (LLM) ensembles. From first principles, we derive an exact decomposition of LLM ensemble lift into rescue and damage masses, which yields a compact heuristic for calculating uplift. From this we extract the metrics which predict ensemble performance: an accuracy-adjusted correctness correlation, $\phi_{\mathrm{adj}}$, together with the accuracy gap and collective accuracy of the pair. We test the law on 767,520 inferences from ten open-weight models over two graduate-level science benchmarks, together with a novel agentic cybersecurity benchmark in which each model conducts digital-forensics inves
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית