Table of Contents What are Multi-Vector Models? The MaxSim Operator What You Gain, and What It Costs Installation Loading a Model Inspecting What a Checkpoint Configured Encoding Queries and Documents Scoring with MaxSim Score Magnitude and MeanMaxSim Semantic Search Retrieve and Rerank Indexing Visual Document Retrieval Audio Retrieval Video Retrieval Interpretability Token Pooling Speeding Up Inference Evaluating a Model Coming from PyLate or colpali-engine Supported Models Text Retrieval Models Visual Document Retrieval Models Acknowledgements Additional Resources Documentation Example Scripts Training Hugging Face Hub Companion Blogposts Sentence Transformers is a Python library for using and training embedding and reranker models for applications like retrieval augmented generation, semantic search, and more. With the v6.0 update, it gains a fourth model type: MultiVectorEncoder, for ColBERT-style late interaction retrieval. Any PyLate checkpoint and any Stanford-NLP ColBERT checkpoint loads straight into it, and colpali-engine models for visual document retrieval can be used too, through the same familiar API you already use for dense, sparse, and reranker models.
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
We’re on a journey to advance and democratize artificial intelligence through open source and open science.

Key points
- Where a regular embedding model compresses a whole text into one vector, a multi-vector model keeps one vector per token and scores query against document with the MaxSim operator.
- If you want to learn how to train them, see the companion Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers blogpost.
- A dense embedding model reads a text and returns a single fixed-size vector.
Sentences selected automatically from the original article by Hugging Face Blog.
Story details
- Published
- By
- Tom Aarsen, Antoine Chaffin, Raphael Sourty
- Format
- Article
- Original
- huggingface.co ↗



