Tencent has released EVIE-Preview-4.5B, a state-of-the-art multilingual Visual Document Retrieval (VDR) model built upon Qwen3.5-4B. The model employs ColBERT-style late interaction with native 128-dimensional multi-vector token embeddings to achieve top-tier performance across ViDoRe benchmarks.

  • Outperforms larger 8B models on ViDoRe V3, leading in 7 out of 8 public domains.
  • Delivers an average accuracy of 85.93 nDCG@5 on ViDoRe V1+V2.
  • Supports robust zero-shot generalization across diverse languages (EN, FR, DE, IT, ES, PT, ZH) and visual formats.
  • Fully integrated with the standard colpali-engine ecosystem for late-interaction scoring pipelines.

The model is trained on approximately 0.8 million high-quality image-query pairs and utilizes dynamic mining and verification to refine training data quality.