יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אין הסכמה אמיתית על רגש: בני אדם, כלים מותאמים ו-LLM נאבקים עם טקסטים ברשתות חברתיות

Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts
מחקר מעריך כלים לניתוח רגש ומודלים LLM, כולל Qwen ו-Llama, ומוצא חוסר הסכמה בין בני אדם וכלים. המודל Twitter-roBERTa-base הראה התאמה חזקה עם דירוגי בני אדם.
תקציר מקורי באנגליתarXiv:2610.10318v1 Announce Type: new Abstract: Social media is a rich source of real-time public sentiment, but widely used sentiment analysis tools are often applied without understanding their limitations. In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets. We measured agreement using two statistical measures: Cohen's kappa for pairwise comparisons and Fleiss' kappa for multiple raters. Even among the human raters, our results showed only fair agreement, highlighting the subjectivity of sentiment analysis. Higher agreement was observed under the binary sentiment classific
קרא במקור המקורי