Researchers introduce an agentic forecasting environment and dataset derived from over 2,100 resolved Polymarket questions to train language models to gather their own evidence. The system allows agents to acquire context via web search and financial time series at rollout time, constrained by leak filtering to ensure information precedes the question's cutoff.
- Trained on Qwen3.5-35B-A3B using single-epoch GRPO with a Brier-score reward.
- Calibration improves by 30-40% and search attempts decrease from 3.8 to 2.25 per rollout as evidence discipline is learned.
- The model outperforms four frontier models, including Claude Opus 4.5 (soft-Brier 0.254 vs. 0.256), at approximately 5% of the inference cost.
The authors release the environment, dataset, and per-rollout records as a reusable harness for temporal forecasting agents.