יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מודל במצוקה: בדיקת תחושות על תפוצה סינתטית של מדיה חברתית

Model in Distress: Sentiment Analysis on French Synthetic Social Media
במאמר זה, פותחים פיפלינג לייצור נתונים סינתטיים כללי, שמשמש לבדיקת תחושות בתפוצה סינתטית של מדיה חברתית בצרפתית. הפיפלינג משתמש בתרגום חזרה עם מודלים שהוטמעו כדי לייצר 1.7 מיליון תפוצות סינתטיות מקורפוס זרע קטן, ומשלב ראיונות סינתטיים. המודלים שהוטמעו, עם פרמטרים של 600 מיליון, השיגו 77-79% דיוק בנתוני בדיקה שנערכו על ידי בני אדם, והגיעו לדיוק של SOTA LLMs ומסמנים מיוחדים.
תקציר מקורי באנגליתarXiv:2604.18226v2 Announce Type: replace Abstract: Automated analysis of customer feedback on social media is hindered by three challenges: the high cost of annotated training data, the scarcity of evaluation sets, especially in multilingual settings, and privacy concerns that prevent data sharing and reproducibility. We address these issues by developing a generalizable synthetic data generation pipeline applied to a case study on customer distress detection in French public transportation. Our approach utilizes backtranslation with fine-tuned models to generate 1.7 million synthetic tweets from a small seed corpus, complemented by synthetic reasoning traces. We train 600M-parameter reasoners with English and French reasoning that achieve 77-79% accuracy on human-annotated evaluation dat
קרא במקור המקורי