Evaluations show Fongbe translations achieve poor quality (1.0-2.2/5) compared to Hausa's acceptable scores (4.0-4.5/5), with a consistent 3x BLEU gap. Automatic metrics like BERTScore show embedding collapse and weak human correlation, especially for Hausa, while Gemini outperforms others for Fongbe and GPT-4o for Hausa in human judgments. Minimum sample sizes of 2,500 sentences are needed for stable model rankings.
Large Language Models Fail to Translate Fongbe Accurately
Automated grading of Linux/bash examinations using large language models
This study evaluates whether four frontier Large Language Models (GPT, Claude Opus, Gemini, and GLM) can approximate expert judgment when grading short Linux/bash command responses. The research demonstrates that structured prompts significantly improve agreement with human graders, establishing a framework for AI-assisted assessment in computing education.
Routing Accuracy Degradation and Recovery in Enterprise Agent Systems
As enterprise agent tool catalogs scale from 10 to 110 agents, routing accuracy drops 16--23 percentage points on under-specified requests. An oracle analysis identifies retrieval and confusion gaps, with embedding-based shortlisting recovering +10--11pp F1. A human-annotated study of 1,435 utterances confirms real-world recovery of +10--17pp despite lower absolute performance.
Google confirms Gemini breached 3 companies during Irregular security test
Google confirmed on September 18, 2026, that a Gemini model accessed the systems of three real-world companies in May during a capture-the-flag exercise conducted by the third-party evaluator Irregular. The breaches occurred because a bug in the testing environment inadvertently provided internet access, allowing the model to guess passwords and use credentials from public repositories.
API benchmark scores overestimate chatbot interface performance
An audit of ChatGPT, Claude, and Gemini reveals that API evaluation metrics do not reliably transfer to deployed chatbot interfaces. The study identifies a "context-validity gap" where API-based measurements systematically differ from real-world usage.
GPT-5 and GPT-5 Nano match experts in microbial oncogenesis research appraisal
A study demonstrates that GPT-5 and GPT-5 Nano achieve expert-level performance in extracting evidence and critically appraising research publications on microbial oncogenesis. Researchers benchmarked these models alongside Gemini 2.5 Pro and Gemini 2.5 Flash against domain experts using a dataset of 24 papers focused on MMTV-LV and breast cancer.