A study of 3,750 queries across five industries finds moderate recommendation concentration, with a mean Gini coefficient of 0.28. Cross-model agreement on top-recommended brands was only 41.6%, and displacement scores varied by industry, ranging from 0.4:1 to 4.3: 1. The results challenge the 'winner-takes-all' narrative and introduce three reproducible metrics for competitive-intelligence analysis.
AI Recommendation Ownership: Empirical Map of Brand Category Ownership
AI-Constructed Brand Reputation Is Language-Bound
AI-generated brand reputations vary significantly by language, with Uralic and Baltic languages showing more positive sentiment and Germanic languages, including English, being more critical. Query language impacts which brands are recommended, especially for local champions, where home-language queries increase visibility by 0.80 points compared to English queries. English-only monitoring fails to capture the full AI visibility of locally headquartered brands, creating a measurable language blind spot.
Automated grading of Linux/bash examinations using large language models
This study evaluates whether four frontier Large Language Models (GPT, Claude Opus, Gemini, and GLM) can approximate expert judgment when grading short Linux/bash command responses. The research demonstrates that structured prompts significantly improve agreement with human graders, establishing a framework for AI-assisted assessment in computing education.
OpenBioRQ: Benchmark for Agentic Biomedical Research Faithfulness
OpenBioRQ introduces a benchmark of 12,553 unsolved biomedical research questions across 12 domains, designed to test agentic models' faithfulness and abstention. It evaluates models in a tool-using setting without answer keys, using real follow-up evidence rather than parametric knowledge, and reveals significant agentic collapse on the hardest questions where tools are no longer used despite being critical.
Large Language Models Fail to Translate Fongbe Accurately
Evaluations show Fongbe translations achieve poor quality (1.0-2.2/5) compared to Hausa's acceptable scores (4.0-4.5/5), with a consistent 3x BLEU gap. Automatic metrics like BERTScore show embedding collapse and weak human correlation, especially for Hausa, while Gemini outperforms others for Fongbe and GPT-4o for Hausa in human judgments. Minimum sample sizes of 2,500 sentences are needed for stable model rankings.
LLM Alignment Using Implicit User Feedback
A new dataset, IFLLM, collects mouse trajectories and eye gazing data from users interacting with LLMs. It shows that implicit feedback significantly improves LLM alignment, boosting text-based reward model accuracy from 55% to 64% and nearly tripling response quality improvements after DPO training on eight LLMs.