Instructions to use nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
Paper: NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval | Blog
NanoVDR-Q-ModernBERT-Qwen3VL2B-2048
⚠️ This model was renamed
It was published as
nanovdr/NanoVDR-L. Old links andfrom_pretrainedcalls still resolve through a redirect, but please move to the new id.- SentenceTransformer("nanovdr/NanoVDR-L") + SentenceTransformer("nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048")Why. NanoVDR started as one thing: a small text encoder that replaces the query side of a large vision-language retriever. It has since grown a second tower that replaces the document side, and a second teacher, so a name like
-Sno longer says enough. Two checkpoints only work together when they were distilled from the same teacher into the same width, and neither fact was recoverable from the old names.The scheme is now
NanoVDR-<Q|D>-<variant>-<teacher>-<width>[-ML], and the rule is simply that the teacher and the width have to match. This model isQ(query tower),ModernBERT(backbone), distilled from Qwen3-VL-Embedding-2B into 2048 dimensions.-MLused to be spelled-Multi, which read as multi-vector when it meant multilingual.
ModernBERT-base ablation variant. For production use, we recommend NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML.
NanoVDR-Q-ModernBERT-Qwen3VL2B-2048 is a 151M-parameter text-only query encoder for visual document retrieval, trained via asymmetric cross-modal distillation from Qwen3-VL-Embedding-2B. It uses ModernBERT-base + a 2-layer MLP projector and achieves the highest v1 score (82.4) among all NanoVDR variants.
Highlights
- Single-vector retrieval — queries and documents share the same 2048-dim embedding space as Qwen3-VL-Embedding-2B; retrieval is a plain dot product, FAISS-compatible, 4 KB per page (float16)
- Lightweight on storage — 612 MB model; doc index costs 64× less than ColPali's multi-vector patches
- Asymmetric setup — tiny 151M text encoder at query time; large VLM indexes documents offline once
Results
| Model | Params | ViDoRe v1 | ViDoRe v2 | ViDoRe v3 | Avg Retention |
|---|---|---|---|---|---|
| Qwen3-VL-Emb (Teacher) | 2.0B | 84.3 | 65.3 | 50.0 | — |
| NanoVDR-Q-ModernBERT-Qwen3VL2B-2048 | 151M | 82.4 | 61.5 | 44.2 | 93.4% |
| NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML | 69M | 82.2 | 61.9 | 46.5 | 95.1% |
NDCG@5 (×100). Retention = Student / Teacher averaged across v1/v2/v3.
Usage
Prerequisite: Documents must be indexed offline using Qwen3-VL-Embedding-2B (the teacher model). See the NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML model page for a complete indexing guide.
from sentence_transformers import SentenceTransformer
import numpy as np
# doc_embeddings: (N, 2048) from teacher indexing (see prerequisite above)
model = SentenceTransformer("nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048")
query_embeddings = model.encode(["What was the revenue growth in Q3?"]) # (1, 2048)
scores = query_embeddings @ doc_embeddings.T
top_k_indices = np.argsort(scores[0])[-5:][::-1]
Training Details
| Value | |
|---|---|
| Architecture | ModernBERT-base (149M) + MLP projector (768 → 768 → 2048, 2.4M) = 151M |
| Objective | Pointwise cosine alignment with teacher query embeddings |
| Data | 711K query-document pairs |
| Epochs / lr | 20 / 2e-4 |
| Training cost | ~11.7 GPU-hours (1× H200) |
| CPU query latency | 109 ms |
All NanoVDR Models
| Model | Backbone | Params | v1 | v2 | v3 | Retention |
|---|---|---|---|---|---|---|
| NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML | DistilBERT | 69M | 82.2 | 61.9 | 46.5 | 95.1% |
| NanoVDR-Q-DistilBERT-Qwen3VL2B-2048 | DistilBERT | 69M | 82.2 | 60.5 | 43.5 | 92.4% |
| NanoVDR-Q-BERT-Qwen3VL2B-2048 | BERT-base | 112M | 82.1 | 62.2 | 44.7 | 94.0% |
| NanoVDR-Q-ModernBERT-Qwen3VL2B-2048 | ModernBERT | 151M | 82.4 | 61.5 | 44.2 | 93.4% |
Want a system with no teacher at all?
Every model in this table targets Qwen3-VL-Embedding-2B at 2048 dimensions, so documents still have to be indexed by that 2B teacher. A second family targets the 8B teacher at 4096 dimensions and includes a document tower, so nothing multi-billion runs at indexing time either:
| Model | Role | Params |
|---|---|---|
| NanoVDR-Q-DistilBERT-Qwen3VL8B-4096-ML | query tower | 70M |
| NanoVDR-D-HiRes-Qwen3VL8B-4096 | document tower | 457M |
| NanoVDR-D-Fast-Qwen3VL8B-4096 | document tower, 3x fewer visual tokens | 457M |
The two families are not interchangeable: 2048-d and 4096-d vectors live in different spaces and will not score against each other.
Citation
@article{nanovdr2026,
title={NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval},
author={Liu, Zhuchenyang and Zhang, Yao and Xiao, Yu},
journal={arXiv preprint arXiv:2603.12824},
year={2026}
}
License
Apache 2.0
- Downloads last month
- 71
Model tree for nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048
Base model
answerdotai/ModernBERT-baseDatasets used to train nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048
openbmb/VisRAG-Ret-Train-Synthetic-data
llamaindex/vdr-multilingual-train
Paper for nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048
Evaluation results
- NDCG@5 on ViDoRe v1self-reported82.400
- NDCG@5 on ViDoRe v2self-reported61.500