Topic · Benchmark results
lab NVIDIA Research · 11d ago · 16 views

ReasoningBomb induces pathologically long reasoning in large reasoning models

Researchers have formalized inference-time denial-of-service (PI-DoS) attacks that exploit the high computational cost of explicit multi-step reasoning in Large Reasoning Models (LRMs). They present ReasoningBomb, a reinforcement-learning-based framework that trains an attacker to generate short natural prompts that drive victim models into pathologically long and often non-terminating reasoning traces.

media MarkTechPost · 21h ago GDPval-AA (Elo) · 1695.0Elo · 13 views

SpaceXAI releases Grok 4.7 with larger base model and improved benchmarks

SpaceXAI has released Grok 4.7, a new flagship model for coding, agentic tasks, and knowledge work that utilizes a larger base model and extended reinforcement learning run compared to Grok 4.6. The model maintains the same pricing structure as its predecessor while offering enhanced capabilities in self-verification, long-context handling, and safety.

lab Hugging Face Blog · 15h ago · 9 views

AISI releases verified evaluation results for six frontier models on five benchmarks

The UK AI Security Institute (AISI) has made publicly reported evaluation methods and findings available through EvalEval's Evaluation Cards platform to improve benchmark reproducibility. This release includes verified results, context, and configuration information for five specific benchmarks: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0.

media MarkTechPost · 7d ago LMSYS Arena (Elo) · 1866.0Elo · 20 views

Prior Labs releases TabPFN-3.5, beating 2015 Otto Kaggle winner with default settings

Prior Labs has released TabPFN-3.5, a tabular foundation model that predicts on a table in a single forward pass without per-dataset training or tuning. The model scores 0.375 on the private leaderboard of the 2015 Otto Group Product Classification Challenge, beating the original winning score of 0.382 achieved by Gilberto Titericz and Stanislav Semenov.