Skip to content
worth noting New models

Hugging Face released open-source multimodal encoders NeoMME for visual document retrieval

clearly official source

Hugging Face released open-source multimodal encoders NeoMME (260M and 800M parameters, Apache 2.0) for document retrieval from page images — according to the company, 2× faster than ColModernVBERT and with an index reduced in size by 255×.

Hugging Face released NeoMME, a family of multimodal encoders in two sizes (260M and 800M parameters), designed to process both text and images with a single bidirectional Transformer. Unlike typical generative vision-language models, NeoMME does not use a separate pretrained vision branch or a causal language model — the entire model was trained from scratch using a masked discrete diffusion objective. The model supports dynamic image resolution, a long context of 16 384 tokens (enough for two 4K UHD pages), and is multilingual thanks to a BPE tokenizer with a vocabulary of 131 thousand tokens.

The company fine-tuned the model for visual document retrieval using the ColPali method (processing full page images instead of text extracted by OCR), as NeoMME-Retriever, with both dense and late-interaction embeddings produced in a single pass. According to the company, both models lie on the Pareto frontier of the ViDoRe v3 benchmark for quality (nDCG@10) versus model size. On an NVIDIA L40S GPU, the 260M variant can process approximately 51 pages per second at a resolution of 2048×2048, which, according to the company, is roughly twice the throughput of ColModernVBERT.

The combination of hierarchical token pooling and asymmetric quantization reduces the late-interaction index size from approximately 1.5 MB to 6 kB per page (a 255-fold reduction), while retaining more than 95 % of the original quality (nDCG@10), according to the company. NeoMME models were pretrained on 524 billion tokens (including 290 billion text-only tokens) using the NorMuon optimizer — a significantly smaller training budget than the 2 trillion tokens used for ModernBERT.

All model checkpoints are available in the Hugging Face Transformers library under the Apache 2.0 license. The full source article was not available; its final section is missing.

What changed

Why it matters

The model processes entire scanned document pages as images and bypasses the need for OCR, preserving layouts, tables and charts — this is directly applicable to deploying RAG and search systems over extensive archives of PDFs and scans. A significantly smaller index (according to the company, 255× smaller while retaining most of the quality) means lower storage costs and faster retrieval when operating at scale, and the Apache 2.0 license allows commercial deployment without licensing restrictions.

Release card

NeoMME

Hcompany

open weights
Specifications
260M and 800M parameters
Context
16 384 tokens (up to two 3840×2160 4K UHD images)
Inputs
text and image inputs (document page images) processed by a single bidirectional Transformer, with vector… as output
Licence
Apache 2.0
Availability
Model checkpoints are available in the Hugging Face Transformers library.
Documented measurements
According to the sources, it is suitable for
  • visual document retrieval directly from page images without OCR (the ColPali method)
  • processing long multilingual documents of up to 16 384 tokens, including text, code and mathematics
  • compact storage of the search index through hierarchical pooling and quantization
Documented limits
  • does not generate text autoregressively; it is only an encoder that produces vector embeddings, not a conversational model

The source compares NeoMME with generative vision-language models that use a separately pretrained vision branch and a causal language decoder, and with ColModernVBERT in terms of page processing speed.

The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →

Two audiences, two different impacts

What this means

01

For individuals

Developers working on visual document retrieval (RAG over PDFs and scans) get an open-source model that, according to the manufacturer, can process pages without OCR and with a significantly smaller index, allowing deployment without having to train their own solution.

What to do Try NeoMME-Retriever in Hugging Face Transformers on a sample document set and compare throughput and index size with the existing solution.
More practical updates →
02

For a business

Companies running search across large document collections (e.g. scanned PDFs, invoices, contracts) may be able to reduce the storage and computing costs associated with retrieval infrastructure through a significantly smaller index and higher throughput, without having to build their own model from scratch.

Development
What to decide Consider testing NeoMME-Retriever as an alternative to the current OCR+embedding pipeline for document search if the company runs retrieval over large numbers of scanned pages.
More business impacts →
Apache 2.0 Hugging Face multimodal encoder NeoMME open-source visual retrieval

Check the original

Event sources

clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
Hugging Face Blog primary source · first detected NeoMME: an efficient Multimodal-native and Multilingual Encoder