כתבה
arXiv cs.AI ·
איך למדוד נוכחות ברחבי ניסויים
Calibrate Globally, Measure Everywhere: Scaling LLM-Based Prevalence Measurement Across A/B Experiments
פינטרסט משתמשת בשיטה חדשה למדידת נוכחות תוכן בניסויים, באמצעות למידת מכונה וכיוונים גלובליים. השיטה מאפשרת מדידה מדויקת יותר ובעלות נמוכה יותר.
תקציר מקורי באנגליתarXiv:2602.16111v2 Announce Type: replace-cross Abstract: Online media platforms track the share of impressions associated with content attributes, or prevalence, to evaluate trade-offs and set guardrails in A/B experiments. LLM-based labeling provides a high-fidelity reference measurement, but is cost-prohibitive to run per experiment, per arm, per segment, and per day on a platform with hundreds of concurrent experiments. We describe a surrogate-based prevalence measurement system deployed in Pinterest's experimentation platform. The contribution is system-level rather than estimator-level: the system maintains a single global calibration of ML score buckets, continuously refreshed from a recurring LLM-labeled stream, and reuses the resulting bucket-level prevalences across every experim
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית