כתבה
arXiv cs.AI ·
QO-Bench: Diagnosing Query-Operator-Preserving Retrieval over Typed Event Tuples
תקציר מקורי באנגליתarXiv:2606.04646v2 Announce Type: replace-cross Abstract: Many real-world questions over business, legal, and scientific corpora are natural-language versions of database-style queries over records latent in text. Existing retrieval-augmented generation (RAG) systems are optimized primarily for semantic relevance, but retrieving plausible passages does not guarantee correct query execution. We introduce QO-Bench, a diagnostic benchmark for query-operator question answering over typed event tuples. The benchmark covers 22,984 news articles and 614 corporate events, with 18 query templates instantiating 785 questions. Each gold answer is deterministically computed from typed event tuples and scored by recall, with answers matched to the gold tuples by exact match rather than an LLM judge. Th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית