A study of 3,750 queries across five industries finds moderate recommendation concentration, with a mean Gini coefficient of 0.28. Cross-model agreement on top-recommended brands was only 41.6%, and displacement scores varied by industry, ranging from 0.4:1 to 4.3: 1. The results challenge the 'winner-takes-all' narrative and introduce three reproducible metrics for competitive-intelligence analysis.
AI Recommendation Ownership: Empirical Map of Brand Category Ownership
AI-Constructed Brand Reputation Is Language-Bound
AI-generated brand reputations vary significantly by language, with Uralic and Baltic languages showing more positive sentiment and Germanic languages, including English, being more critical. Query language impacts which brands are recommended, especially for local champions, where home-language queries increase visibility by 0.80 points compared to English queries. English-only monitoring fails to capture the full AI visibility of locally headquartered brands, creating a measurable language blind spot.
GPT-5 and GPT-5 Nano match experts in microbial oncogenesis research appraisal
A study demonstrates that GPT-5 and GPT-5 Nano achieve expert-level performance in extracting evidence and critically appraising research publications on microbial oncogenesis. Researchers benchmarked these models alongside Gemini 2.5 Pro and Gemini 2.5 Flash against domain experts using a dataset of 24 papers focused on MMTV-LV and breast cancer.
GPT-5 and GPT-5 Nano match experts in microbial oncogenesis evidence extraction
A study benchmarking large language models on systematic evidence synthesis for microbial oncogenesis found that GPT-5 and GPT-5 Nano perform indistinguishably from domain experts. Researchers evaluated Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, and GPT-5 Nano on 24 research papers using a structured template of 77 items across multiple question types.
Automated grading of Linux/bash examinations using large language models
This study evaluates whether four frontier Large Language Models (GPT, Claude Opus, Gemini, and GLM) can approximate expert judgment when grading short Linux/bash command responses. The research demonstrates that structured prompts significantly improve agreement with human graders, establishing a framework for AI-assisted assessment in computing education.
OpenBioRQ: Benchmark for Agentic Biomedical Research Faithfulness
OpenBioRQ introduces a benchmark of 12,553 unsolved biomedical research questions across 12 domains, designed to test agentic models' faithfulness and abstention. It evaluates models in a tool-using setting without answer keys, using real follow-up evidence rather than parametric knowledge, and reveals significant agentic collapse on the hardest questions where tools are no longer used despite being critical.