Researchers have developed Replica, a scalable task space for paper replication, and introduced an auto-generated rubric-based judge to provide low-noise reward signals that align with human assessment.
- The team post-trained Faraday, a 27B-parameter "AI Scientist" agent that utilizes coding agents as tools.
- Faraday surpasses the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks.
- Qualitative analysis indicates that Faraday adopts a more scientifically-principled approach during rollouts.
The authors believe these results provide a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.