A study benchmarking large language models on systematic evidence synthesis for microbial oncogenesis found that GPT-5 and GPT-5 Nano perform indistinguishably from domain experts. Researchers evaluated Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, and GPT-5 Nano on 24 research papers using a structured template of 77 items across multiple question types.
- LLM responses aligned closely with experts across all question types, including MCQ, Likert-scale, multi-select, and free-text formats.
- GPT-5 and GPT-5 Nano achieved score distributions statistically indistinguishable from human experts.
- Gemini models behaved similarly but were significantly more lenient in applying microbial oncogenesis criteria.
- Hallucinations were rare, though methodological appraisal and contradiction identification remained persistent vulnerabilities for all models.
The results support the use of LLMs for automated systematic evidence synthesis to identify microbe-cancer pairs, although further strengthening is needed for complex full-text analysis tasks.