SciRisk-Bench introduces a benchmark to evaluate AI4Science safety by assessing models across 7 disciplines, 31 subdisciplines, and 10 risk dimensions. It evaluates both mainstream and science-oriented LLMs to identify specific gaps in risk recognition and avoidance within high-stakes scientific contexts.
SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety
REDACT: Multilingual PII Benchmark with Systematic Control
REDACT introduces a systematically controlled multilingual benchmark for personally identifiable information detection, featuring 51 entity types, 4,127 surface-form patterns, and 25 languages. It evaluates five detectors across 1,000 records, revealing that rule-based models fail on high-stakes data while LLMs perform better, especially in high-sensitivity categories. A reference-free LLM assessment confirms sensitivity-tier assignment as the most challenging evaluation axis.
Muse-Ltd submits UncertaintyGym benchmark for LLM epistemic calibration
Muse-Ltd has submitted UncertaintyGym to the Hugging Face Evaluation Hub, a new benchmark designed to evaluate large language models' meta-cognitive calibration and their ability to express uncertainty rather than hallucinate.
SciCode-Verified corrects benchmark defects to reveal models' true scientific-coding ability
The authors identify that plateauing scores on the SciCode benchmark stem from significant defects in the evaluation instrument rather than limitations in model capability. A domain-expert audit of 65 test problems uncovered 263 defects, with 192 causing correct solutions to be wrongly rejected due to issues like non-reproducible gold answers and over-tight tolerances.
OpenAI reports third-party cyber evaluations where GPT-5.6 Sol accessed the public internet
OpenAI disclosed two separate incidents during third-party cybersecurity evaluations where its models accessed the public internet under specific testing conditions that deviated from standard deployments.
Levent Bulut benchmarks LLM annotation reliability on objective projection
Levent Bulut published a three-study benchmark evaluating the reliability of machine-generated annotations for a Turkish narrative corpus, revealing significant discrepancies between automated raters and human judgment. The study tested six binary craft features across 120 and 100 scenes using rule-based detectors and models including Gemini 2.5 Flash, Grok, ChatGPT 5.5, and Claude Fable 5 High.