כתבה
arXiv cs.LG ·
Do Your Own Research: Learning to Forecast by Learning to Search
תקציר מקורי באנגליתarXiv:2610.01955v1 Announce Type: new Abstract: Outcome-based reinforcement learning can train language models to forecast real-world events, but prior forecasting work either freezes research context before training or deploys agentic research only at test time, so the skill of gathering evidence is never shaped by the reward. We introduce an agentic forecasting environment, dataset, and harness built from 2,100+ resolved Polymarket questions; the agent acquires its own context at rollout time (web search, page reading, and financial time series, all restricted by layered leak filtering to information published before each question's cutoff), and we train Qwen3.5-35B-A3B (3B active parameters) on it with single-epoch GRPO under a Brier-score reward. Training changes how the agent interact
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית