Cohere has released Embed 5, a new family of embedding models comprising two tiers: Embed 5 Pro for maximum retrieval quality and Embed 5 Fast for lower latency and cost. Both models support text, images, and fused inputs across 100+ languages with a 128K token context window.
- The key architectural feature is that Pro and Fast share a single embedding space, allowing users to index with one model and query with the other without re-indexing.
- Embed 5 supports variable output dimensions (2048 down to 256) and lower-precision outputs (float32, int8, binary), significantly reducing storage requirements for large-scale deployments.
- On ViDoRe V3 benchmarks, Embed 5 Pro averaged 85.8 and Fast averaged 84.5, outperforming Voyage 4 Large (83.7) and Gemini Embedding 2 (83.2).
- The models are generally available via the Cohere API, Model Vault, Microsoft Foundry, and Amazon SageMaker.
The shared embedding space design aims to optimize agentic workloads by allowing high-quality indexing with Pro while maintaining fast query speeds with Fast, reducing latency for agents issuing multiple searches.