Researchers have developed Replica, a scalable task space for paper replication, to address the need for reliable scientific knowledge. They post-trained Faraday, a 27B-parameter "AI Scientist" agent that uses coding agents as tools, to leverage this environment.

  • Replica provides an auto-generated rubric-based judge with low noise that aligns with human assessment of replication quality.
  • Faraday surpasses the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks.
  • Qualitative analysis shows Faraday adopts a more scientifically-principled approach compared to other models.

The results serve as a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.