A new benchmark evaluates AI-generated text-rich images across six domains, including commercial posters and receipts. It reveals significant domain-dependent performance and sensitivity to JPEG compression, highlighting the need for text- and layout-aware detection methods.
Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images
AISI releases verified evaluation results for six frontier models on five benchmarks
The UK AI Security Institute (AISI) has made publicly reported evaluation methods and findings available through EvalEval's Evaluation Cards platform to improve benchmark reproducibility. This release includes verified results, context, and configuration information for five specific benchmarks: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0.
Expert re-grading reveals frontier models near-saturation on physics benchmarks
A study auditing six widely used physics benchmarks finds that reported low scores for frontier language models are largely due to benchmarking errors rather than model deficiencies. Experts reviewed problem statements, reference solutions, and model responses, distinguishing genuine errors from grader mistakes, incorrect references, and ambiguous questions.
Zero-shot evaluation shows Gemini leads LLMs on 13-class emotion taxonomy
A study evaluated three commercial large language models—Claude (claude-sonnet-4-6), ChatGPT (GPT-5.4), and Gemini (gemini-2.5-flash)—on a zero-shot fine-grained emotion classification task using a stratified 1,000-sentence sample from the boltuix/emotions dataset.
Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams
This research investigates the use of large language models to detect scam phone calls in Turkish, a low-resource language where annotated data is scarce. The study introduces the first public multi-modal dataset containing 100 aligned audio-transcript pairs of scam and benign conversations.
AI-Constructed Brand Reputation Is Language-Bound
AI-generated brand reputations vary significantly by language, with Uralic and Baltic languages showing more positive sentiment and Germanic languages, including English, being more critical. Query language impacts which brands are recommended, especially for local champions, where home-language queries increase visibility by 0.80 points compared to English queries. English-only monitoring fails to capture the full AI visibility of locally headquartered brands, creating a measurable language blind spot.