כתבה
arXiv cs.CL ·
Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue
תקציר מקורי באנגליתarXiv:2606.21844v2 Announce Type: replace Abstract: As AI systems integrate into online spaces, differentiating them from humans in conversations is increasingly important. We present Inverse Turing Bench, a benchmark that evaluates LLMs and other models on their ability to differentiate humans and AI in multi-turn text. The benchmark provides a collection of paired dialogue transcripts, wherein one dialogue is between two humans and the other is between a human and an AI. The task is to correctly identify which dialogue is human-only vs. human-AI. We evaluated a preliminary set of models against this benchmark, and found that GPTZero, Claude Opus-4.6, and GPT-5.5 achieve the highest accuracy: 89.41%, 77.92%, and 75.94% respectively. Our results suggest that statistical approaches to detec
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית