Anthropic has published an addendum to the Claude 3 Model Card detailing Claude 3.5 Sonnet, a new model that outperforms Claude 3 Opus in speed and cost while offering enhanced capabilities in coding and visual processing.
- Claude 3.5 Sonnet sets new performance standards on graduate-level science (GPQA), general reasoning (MMLU), and coding proficiency (HumanEval).
- It achieves state-of-the-art results on five vision benchmarks, including MathVista, ChartQA, and DocVQA.
- In an internal agentic coding evaluation, the model solves 64% of problems compared to 38% for Claude 3 Opus.
- Human feedback evaluations show significant win rates over Claude 3 Opus in core capabilities like coding, documents, and creative writing.
- The addendum also reports improved refusal rates on Wildchat and XSTest datasets and near-perfect recall on Needle In A Haystack tests up to 200k tokens.