Multimodal
lab Hugging Face Blog · 20d ago · 34 views

NeoMME releases efficient multimodal encoders with dense and late-interaction retrieval

Researchers introduce NeoMME, a family of 260M and 800M multilingual multimodal encoders that process text and raw image patches using a single bidirectional Transformer. Unlike generative visual language models, NeoMME does not rely on separate pretrained vision towers or causal decoders, training the entire model from scratch with a masked discrete-diffusion objective.