Tencent has released EVIE-Preview-4.5B, a state-of-the-art multilingual Visual Document Retrieval (VDR) model built upon Qwen3.5-4B. The model employs ColBERT-style late interaction with native 128-dimensional multi-vector token embeddings to achieve top-tier performance across ViDoRe benchmarks.
- Outperforms larger 8B models on ViDoRe V3, leading in 7 out of 8 public domains.
- Delivers an average accuracy of 85.93 nDCG@5 on ViDoRe V1+V2.
- Supports robust zero-shot generalization across diverse languages (EN, FR, DE, IT, ES, PT, ZH) and visual formats.
- Fully integrated with the standard colpali-engine ecosystem for late-interaction scoring pipelines.
The model is trained on approximately 0.8 million high-quality image-query pairs and utilizes dynamic mining and verification to refine training data quality.