יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בדיקה ותיקון כישלונות LLM בצינור מידע ל-SQL

Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
חוקרים בדקו את הסכמתם של LLM עם אנוטטורים אנושיים בצינור מידע ל-SQL ומצאו כי Qwen3.6-27B היא חלופה טובה יותר מאשר GPT-4. הם גם גילו ששילוב של שלושה LLM חזקים יכול לשפר את הסכמתם.
תקציר מקורי באנגליתarXiv:2609.30290v1 Announce Type: cross Abstract: Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen's kappa = 0.04 on a disagreement-enriched set and 0.42 on a uniform-random spot-check, over-flagging 77.1% of the human-FAITHFUL cases in the enriched set. Most of its over-flags trace back to a single mechanism we call GRADE-HALLUCINATION. A self-hosted Qwen3.6-27B replacement (kappa = 0.72) lands in the same range as Claude Opus 4.7 (kappa = 0.71); the head-to-head is underpowered at n = 96, but for the deployment decision that hardly matters, since Qwen costs roughly 1/300 as much per call. Ensembling does not
קרא במקור המקורי