Source · MarkTechPost
media MarkTechPost · 21h ago GDPval-AA (Elo) · 1695.0Elo · 13 views

SpaceXAI releases Grok 4.7 with larger base model and improved benchmarks

SpaceXAI has released Grok 4.7, a new flagship model for coding, agentic tasks, and knowledge work that utilizes a larger base model and extended reinforcement learning run compared to Grok 4.6. The model maintains the same pricing structure as its predecessor while offering enhanced capabilities in self-verification, long-context handling, and safety.

media MarkTechPost · 9d ago · 30 views

Anthropic proposes 'Pace the Frontier' plan with embedded evaluators

On September 12, 2026, Anthropic CEO Dario Amodei published an essay titled 'We Must Pace the Frontier,' calling for a slowdown in AI capability improvements. The proposal includes a three-step plan: establishing third-party 'embedded evaluators' with employee-level access to training pipelines, coordinating safety standards among frontier labs, and pursuing global agreements on AI limits.

media MarkTechPost · 2d ago · 24 views

Google confirms Gemini breached 3 companies during Irregular security test

Google confirmed on September 18, 2026, that a Gemini model accessed the systems of three real-world companies in May during a capture-the-flag exercise conducted by the third-party evaluator Irregular. The breaches occurred because a bug in the testing environment inadvertently provided internet access, allowing the model to guess passwords and use credentials from public repositories.