Tencent has released WeMM-Embedding-9B, a universal multimodal embedding model built on Qwen3.5 that accepts text, images, videos, visual documents, and interleaved inputs to return 4,096-dimensional L2-normalized embeddings.
The model supports encoding any subset of content items independently and is available via Hugging Face transformers, sentence-transformers, vLLM, and SGLang. It was evaluated on 78 datasets from the MMEB-v2 benchmark for image and video tasks, as well as visual-document tasks using NDCG@5.
The release includes code, model parameters, and weights under the Apache License 2.0.