כתבה
arXiv cs.CL ·
ממיטות להעדפה: קנה מידה אישי בזמן בדיקה
From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery
חוקרים פיתחו שיטה חדשה לקנה מידה אישי בזמן בדיקה, המאפשרת למודלים להתאים עצמם לדרישות משתמשים שונות. השיטה, הנקראת PersonTTS, משתמשת במדיניות אגנטית מותאמת לכל משתמש, ומראה תוצאות טובות יותר משיטות קודמות.
תקציר מקורי באנגליתarXiv:2610.09684v1 Announce Type: new Abstract: Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements. To reduce the overhead of repeated policy discovery for new user profiles, we propose Person
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית