יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בנצ'מרקים אישיות של מודלי שפה גדולים

Benchmarking the Personalization Capabilities of Large Language Models
SDR-Arena הוא כלי לבנצ'מרק אישיות של מודלי שפה גדולים. המחקר משווה את Claude Sonnet 4.6 ומודלים אחרים ביכולתם לייצר תוכן מותאם. התוצאות מראות כי המודלים מצליחים לשחזר רק חלק מהתוכן המנצח.
תקציר מקורי באנגליתarXiv:2607.20471v2 Announce Type: replace Abstract: Personalization is classically a two-party problem: a sender chooses what to say, and a receiver with independent objectives decides whether to act. A salesperson pitching the same analytics product leads with HIPAA compliance for a hospital and real-time reporting for a retailer, expecting a different argument to work on each. Existing LLM personalization benchmarks measure a narrower, one-party property: whether output matches the preferences of the same user it serves-sender and receiver being the same, as when RLHF aligns an assistant to its own user. The two-party case is harder to study automatically, since it needs ground truth linking specific content to an observed receiver action. Sales outreach provides this: a message written
קרא במקור המקורי