Mistral has released Mistral Large 2, a new large language model featuring 123 billion parameters designed for high-throughput single-node inference. The model supports a 128k context window and dozens of languages while offering enhanced performance in coding, reasoning, and instruction following.
- The pretrained version achieves 84.0% accuracy on MMLU, setting a new point on the performance/cost Pareto front for open models.
- It performs on par with leading models such as GPT-4o, Claude 3 Opus, and Llama 3 405B in code and reasoning tasks.
- The model is trained to minimize hallucinations and acknowledge when it lacks sufficient information for confident answers.
- Weights are available under the Mistral Research License for research and non-commercial use, with commercial self-deployment requiring a separate license.
The release aims to provide a cost-efficient alternative for building innovative AI applications, with availability on la Plateforme, HuggingFace, and major cloud providers including Google Cloud Platform, Azure, Amazon Bedrock, and IBM watsonx.ai.