Topic · Evaluation & benchmarks
lab NVIDIA Research · 12d ago · 8 views

NVIDIA introduces Spatial-IQ, a hierarchical diagnostic framework for multimodal model spatial reasoning

Researchers at NVIDIA have introduced Spatial-IQ, a diagnostic framework designed to deconstruct the spatial intelligence of multimodal large language models (MLLMs). Unlike existing benchmarks that treat models as black boxes, this approach decomposes object counting in stacked 3D structures into nine perceptual and cognitive sub-tasks aligned with human developmental stages.

media Hugging Face Forums · 6d ago · 1 view

GPT-5.6 recovers 95% of hidden messages in unseen literary cryptography benchmark

A working paper by Joseph JM Walker evaluates frontier language models on an unseen multi-channel literary cryptography benchmark embedded in the physical novel "I Wrote a Book and Made a Million Dollars (I.B. Wryten)." The study finds that GPT-5.6 independently recognized the secondary communication channel and recovered approximately 95% of the embedded material on its first pass, whereas earlier models demonstrated effectively 0% ability.

media Hugging Face Forums · 7d ago · 8 views

Levent Bulut benchmarks LLM annotation reliability on objective projection

Levent Bulut published a three-study benchmark evaluating the reliability of machine-generated annotations for a Turkish narrative corpus, revealing significant discrepancies between automated raters and human judgment. The study tested six binary craft features across 120 and 100 scenes using rule-based detectors and models including Gemini 2.5 Flash, Grok, ChatGPT 5.5, and Claude Fable 5 High.

arxiv arXiv cs.CL · 12d ago

Zing framework improves LLM social intelligence via SoMBench benchmark and Actio grounding

The report introduces Zing, an integrated framework designed to enhance the social intelligence of large language models by measuring, internalizing, and grounding social capabilities. The authors present SoMBench, a psychology-grounded benchmark spanning 17 secondary dimensions and 3,481 expert-verified instances, which reveals that current LLMs have significant room for improvement with no model reaching near-ceiling performance.