כתבה
arXiv cs.LG ·
When Can Conformal Risk Control Certify LLM Outputs? Bounds, Impossibility, and Adaptation for Structured Generation
תקציר מקורי באנגליתarXiv:2606.29054v2 Announce Type: replace Abstract: Large language models (LLMs) deployed for structured generation (NER, JSON extraction, QA, and classification) lack formal reliability guarantees, and standard heuristic abstention policies violate user-specified risk targets on 7.5--12.5% of settings, with no per-deployment guarantee. We characterize when conformal risk control (CRC) can certify structured LLM outputs and when it provably cannot. First, we prove a sharpened, attained feasibility frontier: when base risk \mu exceeds target \alpha, any distribution-free method must abstain on at least (\mu-\alpha)/(M-\alpha) of inputs, yielding a closed-form feasibility test that decides whether CRC can work before running it. Second, we establish a proven certification phase diagram acros
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית