A new benchmark evaluates AI-generated text-rich images across six domains, including commercial posters and receipts. It reveals significant domain-dependent performance and sensitivity to JPEG compression, highlighting the need for text- and layout-aware detection methods.
Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images
Zero-shot evaluation shows Gemini leads LLMs on 13-class emotion taxonomy
A study evaluated three commercial large language models—Claude (claude-sonnet-4-6), ChatGPT (GPT-5.4), and Gemini (gemini-2.5-flash)—on a zero-shot fine-grained emotion classification task using a stratified 1,000-sentence sample from the boltuix/emotions dataset.
Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams
This research investigates the use of large language models to detect scam phone calls in Turkish, a low-resource language where annotated data is scarce. The study introduces the first public multi-modal dataset containing 100 aligned audio-transcript pairs of scam and benign conversations.
AI-Constructed Brand Reputation Is Language-Bound
AI-generated brand reputations vary significantly by language, with Uralic and Baltic languages showing more positive sentiment and Germanic languages, including English, being more critical. Query language impacts which brands are recommended, especially for local champions, where home-language queries increase visibility by 0.80 points compared to English queries. English-only monitoring fails to capture the full AI visibility of locally headquartered brands, creating a measurable language blind spot.
Benchmarking Agentic Review Systems for AI-Assisted Research
A study evaluates four AI review systems across six language models, finding OpenAIReview with GPT-5.5 achieves 83.0% accuracy in matching paper quality to external signals and detects 71.6% of injected errors. Real user feedback shows positive sentiment, with a 1.44-to-1 vote ratio, though false positives and minor nitpicks remain common.
REDACT: Multilingual PII Benchmark with Systematic Control
REDACT introduces a systematically controlled multilingual benchmark for personally identifiable information detection, featuring 51 entity types, 4,127 surface-form patterns, and 25 languages. It evaluates five detectors across 1,000 records, revealing that rule-based models fail on high-stakes data while LLMs perform better, especially in high-sensitivity categories. A reference-free LLM assessment confirms sensitivity-tier assignment as the most challenging evaluation axis.