Researchers introduce TasteVal, a benchmark designed to evaluate the experimental research taste of frontier AI models by measuring their ability to design experiments and interpret results efficiently. The study assesses 20 models released between 2023 and 2026 using eight novel tasks where a Researcher agent iteratively designs experiments while a fixed Coder agent implements them.

  • TasteVal operationalizes experimental taste as compute efficiency, defining it as the multiplier of experimental compute relative to human experts.
  • The evaluation compares model performance against a baseline established by 24 recruited human experts across the tasks.
  • Opus 5.5 emerged as the best-performing model, exceeding the expert baseline with a compute multiplier of 2.3x at roughly 1/30th of the baseliners' average cost.
  • The data indicates that the compute multiplier for frontier models has doubled approximately every 3.0 months since December 2025, accelerating from the previous rate of every 14 months.

The benchmark highlights a significant acceleration in AI efficiency gains, suggesting that experimental taste is becoming a critical factor in forecasting future AI progress.