Sentence Transformers v6: ColBERT-Style Multi-Vector Models Fully Supported
Melvin Vivas · X post · 2026-08-20 · Open on X
Topics: Embeddings & Vector Databases, Retrieval-Augmented Generation (RAG) · Level: intermediate
Summary
Sentence Transformers v6.0 has been released. It adds MultiVectorEncoder, which makes ColBERT-style late-interaction models a full model type alongside dense, sparse and reranker models. Training, inference and interpretation are all supported.
Key points
- Sentence Transformers v6.0 adds the MultiVectorEncoder model type.
- ColBERT-style late interaction stores one vector per token instead of one vector per document.
- Multi-vector models can now be trained, run and interpreted inside the library.
- The library now covers four model families: dense, sparse, reranker (cross-encoder) and multi-vector.
Resources mentioned
- Sentence Transformers · tool · sbert.net · free
Open-source Python library for training and running embedding, sparse, reranker and (from v6) multi-vector models. - ColBERT · paper · arxiv.org · free
Late-interaction retrieval model that scores documents by matching their token embeddings against the query's token embeddings.
Try this
- Upgrade to Sentence Transformers v6 and try MultiVectorEncoder for late-interaction retrieval.
- Compare dense, sparse and ColBERT-style multi-vector retrieval on your own RAG dataset.
More in Embeddings & Vector Databases
- Liquid AI LFM 2.5 Encoder: a CPU-friendly encoder model
- How RAG Finds the Right Document Fast: Graph-Based Vector Search
- Building Interactive AI Concept Demos with Claude Design (Embeddings Example)
- Gemini Embedding 2: One Vector Space for Five Modalities
- Fine-Tuning MiniLM Embeddings with Synthetic Data, Running on CPU