A study demonstrates that GPT-5 and GPT-5 Nano achieve expert-level performance in extracting evidence and critically appraising research publications on microbial oncogenesis. Researchers benchmarked these models alongside Gemini 2.5 Pro and Gemini 2.5 Flash against domain experts using a dataset of 24 papers focused on MMTV-LV and breast cancer.

  • The evaluation used a structured template with 77 items across multiple question types, including MCQs, Likert scales, and free-text responses.
  • GPT-5 and GPT-5 Nano produced score distributions indistinguishable from human experts, while Gemini models were significantly more lenient in applying criteria.
  • Hallucinations were rare across all models, but methodological appraisal and contradiction identification within full texts remained persistent vulnerabilities.

The findings support the use of LLMs for automated systematic evidence synthesis, although further strengthening is needed for complex methodological tasks.