Hugging Face released open-source multimodal encoders NeoMME for visual document retrieval
Hugging Face released open-source multimodal encoders NeoMME (260M and 800M parameters, Apache 2.0) for document retrieval from page images — according to the company, 2× faster than ColModernVBERT and with an index reduced in size by 255×.
Hugging Face released NeoMME, a family of multimodal encoders in two sizes (260M and 800M parameters), designed to process both text and images with a single bidirectional Transformer. Unlike typical generative vision-language models, NeoMME does not use a separate pretrained vision branch or a causal language model — the entire model was trained from scratch using a masked discrete diffusion objective. The model supports dynamic image resolution, a long context of 16 384 tokens (enough for two 4K UHD pages), and is multilingual thanks to a BPE tokenizer with a vocabulary of 131 thousand tokens.
The company fine-tuned the model for visual document retrieval using the ColPali method (processing full page images instead of text extracted by OCR), as NeoMME-Retriever, with both dense and late-interaction embeddings produced in a single pass. According to the company, both models lie on the Pareto frontier of the ViDoRe v3 benchmark for quality (nDCG@10) versus model size. On an NVIDIA L40S GPU, the 260M variant can process approximately 51 pages per second at a resolution of 2048×2048, which, according to the company, is roughly twice the throughput of ColModernVBERT.
The combination of hierarchical token pooling and asymmetric quantization reduces the late-interaction index size from approximately 1.5 MB to 6 kB per page (a 255-fold reduction), while retaining more than 95 % of the original quality (nDCG@10), according to the company. NeoMME models were pretrained on 524 billion tokens (including 290 billion text-only tokens) using the NorMuon optimizer — a significantly smaller training budget than the 2 trillion tokens used for ModernBERT.
All model checkpoints are available in the Hugging Face Transformers library under the Apache 2.0 license. The full source article was not available; its final section is missing.
Why it matters
The model processes entire scanned document pages as images and bypasses the need for OCR, preserving layouts, tables and charts — this is directly applicable to deploying RAG and search systems over extensive archives of PDFs and scans. A significantly smaller index (according to the company, 255× smaller while retaining most of the quality) means lower storage costs and faster retrieval when operating at scale, and the Apache 2.0 license allows commercial deployment without licensing restrictions.
Release card
NeoMME
Hcompany
- Specifications
- 260M and 800M parameters
- Context
- 16 384 tokens (up to two 3840×2160 4K UHD images)
- Inputs
- text and image inputs (document page images) processed by a single bidirectional Transformer, with vector… as output
- Licence
- Apache 2.0
- Availability
- Model checkpoints are available in the Hugging Face Transformers library.
- Rychlost enkódování stránek na GPU NVIDIA L40S (2048×2048) přibližně 51 stránek za sekundu (model 260M) Reports the speed at which NeoMME 260M processes document page images, which, according to the source, is roughly twice the throughput of ColModernVBERT.
- visual document retrieval directly from page images without OCR (the ColPali method)
- processing long multilingual documents of up to 16 384 tokens, including text, code and mathematics
- compact storage of the search index through hierarchical pooling and quantization
- does not generate text autoregressively; it is only an encoder that produces vector embeddings, not a conversational model
The source compares NeoMME with generative vision-language models that use a separately pretrained vision branch and a causal language decoder, and with ColModernVBERT in terms of page processing speed.
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →
Two audiences, two different impacts
What this means
For individuals
Developers working on visual document retrieval (RAG over PDFs and scans) get an open-source model that, according to the manufacturer, can process pages without OCR and with a significantly smaller index, allowing deployment without having to train their own solution.
For a business
Companies running search across large document collections (e.g. scanned PDFs, invoices, contracts) may be able to reduce the storage and computing costs associated with retrieval infrastructure through a significantly smaller index and higher throughput, without having to build their own model from scratch.
DevelopmentCheck the original
Event sources
clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.