כתבה
arXiv cs.CL ·
Do LLM Debates Repeat Arguments Differently Across Languages?
תקציר מקורי באנגליתarXiv:2607.23442v1 Announce Type: new Abstract: LLM debate is usually evaluated by final answers, but transcripts also reveal whether later turns develop new argumentative content or return to earlier claims in new wording. We study this process with \textit{prior-argument similarity}, an aggregate diagnostic comparing extracted argument units with earlier units in the same debate. In controlled eight-turn debates over 71 motions, six languages, and four model agents, Chinese is the only tested language with a consistently positive gap relative to English across three multilingual embedding models. The gap persists across agents, turn positions, regression adjustment, metric variants, extraction-length controls, a second-extractor subset, and cross-encoder tail rescoring. Manual calibratio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית