NanoVDR

Paper: NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval | Blog

NanoVDR-Q-ModernBERT-Qwen3VL2B-2048

⚠️ This model was renamed

It was published as nanovdr/NanoVDR-L. Old links and from_pretrained calls still resolve through a redirect, but please move to the new id.

- SentenceTransformer("nanovdr/NanoVDR-L")
+ SentenceTransformer("nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048")

Why. NanoVDR started as one thing: a small text encoder that replaces the query side of a large vision-language retriever. It has since grown a second tower that replaces the document side, and a second teacher, so a name like -S no longer says enough. Two checkpoints only work together when they were distilled from the same teacher into the same width, and neither fact was recoverable from the old names.

The scheme is now NanoVDR-<Q|D>-<variant>-<teacher>-<width>[-ML], and the rule is simply that the teacher and the width have to match. This model is Q (query tower), ModernBERT (backbone), distilled from Qwen3-VL-Embedding-2B into 2048 dimensions. -ML used to be spelled -Multi, which read as multi-vector when it meant multilingual.

ModernBERT-base ablation variant. For production use, we recommend NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML.

NanoVDR-Q-ModernBERT-Qwen3VL2B-2048 is a 151M-parameter text-only query encoder for visual document retrieval, trained via asymmetric cross-modal distillation from Qwen3-VL-Embedding-2B. It uses ModernBERT-base + a 2-layer MLP projector and achieves the highest v1 score (82.4) among all NanoVDR variants.

Highlights

  • Single-vector retrieval — queries and documents share the same 2048-dim embedding space as Qwen3-VL-Embedding-2B; retrieval is a plain dot product, FAISS-compatible, 4 KB per page (float16)
  • Lightweight on storage — 612 MB model; doc index costs 64× less than ColPali's multi-vector patches
  • Asymmetric setup — tiny 151M text encoder at query time; large VLM indexes documents offline once

Results

Model Params ViDoRe v1 ViDoRe v2 ViDoRe v3 Avg Retention
Qwen3-VL-Emb (Teacher) 2.0B 84.3 65.3 50.0
NanoVDR-Q-ModernBERT-Qwen3VL2B-2048 151M 82.4 61.5 44.2 93.4%
NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML 69M 82.2 61.9 46.5 95.1%

NDCG@5 (×100). Retention = Student / Teacher averaged across v1/v2/v3.

Usage

Prerequisite: Documents must be indexed offline using Qwen3-VL-Embedding-2B (the teacher model). See the NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML model page for a complete indexing guide.

from sentence_transformers import SentenceTransformer
import numpy as np

# doc_embeddings: (N, 2048) from teacher indexing (see prerequisite above)

model = SentenceTransformer("nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048")
query_embeddings = model.encode(["What was the revenue growth in Q3?"])  # (1, 2048)

scores = query_embeddings @ doc_embeddings.T
top_k_indices = np.argsort(scores[0])[-5:][::-1]

Training Details

Value
Architecture ModernBERT-base (149M) + MLP projector (768 → 768 → 2048, 2.4M) = 151M
Objective Pointwise cosine alignment with teacher query embeddings
Data 711K query-document pairs
Epochs / lr 20 / 2e-4
Training cost ~11.7 GPU-hours (1× H200)
CPU query latency 109 ms

All NanoVDR Models

Model Backbone Params v1 v2 v3 Retention
NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML DistilBERT 69M 82.2 61.9 46.5 95.1%
NanoVDR-Q-DistilBERT-Qwen3VL2B-2048 DistilBERT 69M 82.2 60.5 43.5 92.4%
NanoVDR-Q-BERT-Qwen3VL2B-2048 BERT-base 112M 82.1 62.2 44.7 94.0%
NanoVDR-Q-ModernBERT-Qwen3VL2B-2048 ModernBERT 151M 82.4 61.5 44.2 93.4%

Want a system with no teacher at all?

Every model in this table targets Qwen3-VL-Embedding-2B at 2048 dimensions, so documents still have to be indexed by that 2B teacher. A second family targets the 8B teacher at 4096 dimensions and includes a document tower, so nothing multi-billion runs at indexing time either:

Model Role Params
NanoVDR-Q-DistilBERT-Qwen3VL8B-4096-ML query tower 70M
NanoVDR-D-HiRes-Qwen3VL8B-4096 document tower 457M
NanoVDR-D-Fast-Qwen3VL8B-4096 document tower, 3x fewer visual tokens 457M

The two families are not interchangeable: 2048-d and 4096-d vectors live in different spaces and will not score against each other.

Citation

@article{nanovdr2026,
  title={NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval},
  author={Liu, Zhuchenyang and Zhang, Yao and Xiao, Yu},
  journal={arXiv preprint arXiv:2603.12824},
  year={2026}
}

License

Apache 2.0

Downloads last month
71
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048

Finetuned
(1402)
this model

Datasets used to train nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048

Paper for nanovdr/NanoVDR-Q-ModernBERT-Qwen3VL2B-2048

Evaluation results