Researchers have developed Replica, a scalable task space for paper replication, to address the need for reliable scientific knowledge. They post-trained Faraday, a 27B-parameter "AI Scientist" agent that uses coding agents as tools, to leverage this environment.
- Replica provides an auto-generated rubric-based judge with low noise that aligns with human assessment of replication quality.
- Faraday surpasses the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks.
- Qualitative analysis shows Faraday adopts a more scientifically-principled approach compared to other models.
The results serve as a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.