Anthropic has introduced Claude 3.5 Haiku and an upgraded version of Claude 3.5 Sonnet, advancing capabilities in reasoning, coding, and visual processing. The most significant addition to the upgraded Sonnet model is its ability to interpret GUI screenshots and generate tool calls to perform tasks autonomously.
- Upgraded Claude 3.5 Sonnet achieves a state-of-the-art average success rate of 14.9% on the OSWorld benchmark using only screenshot inputs, rising to 22% with increased interaction steps.
- The new Claude 3.5 Haiku model demonstrates strong performance in reasoning and instruction following, often comparable to original Claude 3.5 Sonnet and Claude 3 Opus.
- Both models underwent extensive safety evaluations, including multimodal red-teaming for computer use risks, and joint pre-deployment testing by the US and UK AI Safety Institutes.
- The knowledge cutoff for the upgraded Claude 3.5 Sonnet is April 2024, while Claude 3.5 Haiku has a cutoff of July 2024.
These updates provide users with enhanced automation capabilities through computer use and improved agentic task completion, while maintaining rigorous safety standards aligned with Anthropic's Responsible Scaling Policy.