Table of Contents What are Multi-Vector models? Why Finetune? Training Components Model Finetuning an existing multi-vector model Building one from a base transformer Which starting point should you pick? Dataset Data on the Hugging Face Hub Local Data Dataset Format Loss Function Training Arguments Evaluator Trainer Callbacks Multi-Dataset Training Evaluation Optimizing the index Acknowledgements Additional Resources Training Examples Documentation Sentence Transformers is a Python library for using and training embedding and reranker models for a wide range of applications, such as retrieval augmented generation, semantic search, semantic textual similarity, and more. Its v6.0 update introduces a fourth model type: MultiVectorEncoder, for ColBERT-style late interaction retrieval, alongside a complete training approach for it. In this blogpost, I'll show you how to use it to finetune a multi-vector model that outperforms general-purpose retrievers on your data. This method can also train strong new multi-vector models from scratch. Everything below runs on pip install -U "sentence-transformers[train]".