On October 7, 2026, Perplexity published "Multimodal embeddings beyond a single vector," introducing pplx-embed-v2-late: a family of embedding models that stop squeezing each document into one vector. It follows the contextual model pplx-embed-v2-context that Perplexity previewed about a week earlier and continues a generation of embedding releases that began with pplx-embed-v1 in February.
This post explains what changed, how late interaction works in plain terms, what it costs in storage, how it compares with the single-vector approach, and what remains unverified. A sourcing note: Perplexity's own page was not reachable from our research environment, so we rely on search summaries of the post and third-party coverage. We could not find the specific benchmark scores, license terms or pricing, and we say so below.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What was released? | pplx-embed-v2-late, a family of late-interaction embedding models |
| What is new? | Token-level vectors instead of one vector per input |
| Which inputs? | Text and images, including rendered PDF pages without OCR |
| Which sizes? | 0.6B and 9B, sharing one embedding space |
| Why share a space? | Index with the 9B model, query with the 0.6B model |
| Main cost? | A larger index: many vectors per document |
| Are the benchmarks confirmed? | Perplexity claims state of the art; we could not verify the scores |
| License and pricing? | Not confirmed in the sources we could reach |
The problem: one vector per document is a bottleneck
A standard embedding model reads a passage and outputs a single list of numbers, say 1,024 of them. Search then compares your query vector to every document vector and returns the closest. That is fast and simple, and it powers most retrieval-augmented generation today.
The compression is lossy by design. A document with five ideas must be squeezed into the same-size vector as a document with one. A query that has two parts, for example "revenue growth in Europe compared with Asia," has to match a single point that blends everything the page says. Research supports the concern: a 2025 paper argued that retrievers returning a single embedding cannot capture targets spread across several regions of the space, and a 2026 paper reported that multi-vector embeddings are provably more expressive than single-vector ones.
What late interaction does instead
Late interaction keeps one small vector per token, or per image patch, and delays the comparison until query time. The scoring step, widely called MaxSim after the ColBERT family of models, works like this:
- Embed the query into token vectors.
- Embed each document into token vectors, ahead of time, and store them.
- For each query token, find the document token it matches best.
- Add up those best-match scores to get the document's score.
The effect is that different parts of the query can latch onto different parts of the document. The word "Europe" can match one sentence, "Asia" another, and "growth" a third, with no need for one vector to represent all of it. Perplexity describes exactly this benefit: "token-level vectors for finer-grained query-document comparisons."
If you want to try the technique on text with open tooling, our coverage of Sentence Transformers v6 and multi-vector ColBERT-style retrieval shows the library support, and our guide to semantic, vector and hybrid search places it among the other retrieval strategies.
The multimodal part: PDF pages without OCR
The headline feature is that the models take images. Perplexity says they can do text-to-image retrieval and search rendered PDF pages directly, without OCR or parsed text.
Why that matters: the usual pipeline for a PDF is to run OCR or a parser, chunk the text, and embed it. That flattens charts, tables, diagrams, multi-column layouts and handwriting, and any parsing error is baked into the index. If the model sees the rendered page as an image and keeps patch-level vectors, it can match a query to the part of the page that contains the chart or table, in context. The family of approaches is often called visual document retrieval, and benchmarks such as ViDoRe measure it. Our coverage of Cohere Embed 5 on ViDoRe v3 and of Cohere Parse 5 and document parsing shows the alternative that still parses text.
One embedding space, two sizes
Perplexity released 0.6B and 9B models and says both share one embedding space. Practically, that lets you:
- Index with the 9B model once, when you can afford the compute, to get the highest-quality document vectors.
- Query with the 0.6B model, which is cheap and fast, so each search stays low-latency.
This is the same pattern Cohere shipped with Embed 5, where you index with Pro and query with Fast in one space, as covered in our Cohere Embed 5 post. It is an answer to a real production problem: indexing is a one-time cost, queries are a recurring one. The caveat is that mixing models only works if they were trained to align, so do not assume it holds for models from different families or versions.
The trade-off: storage and compute
Late interaction has a reputation for being storage and compute hungry, because you keep many vectors per document instead of one. Here is an illustrative calculation. These numbers are our own assumptions, not Perplexity's, since we do not know its per-token vector dimension.
| Setup | Vectors per 1,000-token document | Dimension | Bytes per value | Storage per document |
|---|---|---|---|---|
| Single vector | 1 | 1,024 | 1 (int8) | about 1 KB |
| Multi-vector, one per token | 1,000 | 128 | 1 (int8) | about 128 KB |
Under those assumptions the multi-vector index is roughly 100 times larger. A million documents would move from about 1 GB to about 128 GB. Real systems reduce this with smaller dimensions, quantization, pruning and approximate search, and images add patch vectors on top. The point is not the exact figure, but that the index size scales with tokens, so plan for it. Search also gets more complex, because you typically retrieve candidates with a cheaper method and then rescore them with MaxSim.
For the cost side of single-vector systems, see our write-up of Perplexity's fast embedding serving infrastructure and the Google TurboVec vector search work.
How it differs from pplx-embed-v2-context
Perplexity now has two v2 directions, and they solve different problems.
| pplx-embed-v2-context | pplx-embed-v2-late | |
|---|---|---|
| Vectors per item | One per chunk | Many, at token or patch level |
| Main idea | Encode each chunk knowing its whole document | Match parts of a query to parts of a document |
| Extra inference cost | None reported at query time | Higher storage and scoring cost |
| Modalities | Text | Text and images |
| Best for | Chunked documents where context matters | Multi-part queries, visual documents, fine-grained matching |
Our contextual model post covers the first in detail. You might use both: contextual embeddings for cheap first-pass retrieval and late interaction to rerank.
What is confirmed and what is not
| Item | Status |
|---|---|
| Name, family and release date (October 7, 2026) | Reported by Perplexity's post and coverage |
| Token-level vectors, text and image input | Reported |
| PDF pages searched without OCR | Reported by Perplexity |
| 0.6B and 9B sizes in one shared space | Reported |
| State-of-the-art on vision and text retrieval benchmarks | Perplexity's claim; specific scores not confirmed by us |
| 53% on BrowseComp-Plus for a 0.6B late model | A social post attributed to a Perplexity researcher; unverified and not a ViDoRe or MTEB score |
| License, weights on Hugging Face, API pricing | Not confirmed; older v1 pricing on aggregators is not a guide |
Be careful with aggregator pages: some list prices and licenses for earlier pplx-embed models that may not apply here. Check Perplexity's official model cards and API docs before building on it.
What this means for what you build
- If your documents are visual, such as slide decks, scanned reports, invoices and charts, test page-image retrieval against your OCR pipeline on 50 real queries. Measure recall and the cost of the index.
- If your queries are multi-part, late interaction is the technique to try, text-only or multimodal.
- Budget for the index. Estimate vectors per document times dimension times bytes, then decide on quantization and a two-stage search.
- Use the shared space. If you adopt the 9B for indexing, test whether the 0.6B query model keeps quality at your latency target.
- Keep your evals. Vendor benchmarks are not your corpus. Our guide to embedding models lists alternatives to compare against, and Mixedbread's search agent shows another direction for retrieval stacks.
Related reading on explainx.ai
- Perplexity pplx-embed-v2-context 9B preview
- Cohere Embed 5: Pro and Fast on ViDoRe v3
- Sentence Transformers v6 and multi-vector ColBERT retrieval
- What are embeddings? Vector search guide
- Semantic vs vector vs hybrid search
- Top 10 open and closed embedding models
- Perplexity fast embeddings and GPU serving
- Perplexity Search API and index debut
Details come from Perplexity's October 7, 2026 post as summarized in search results and third-party coverage; we did not read the primary page or run the models. The storage table is our own illustrative arithmetic. Check Perplexity's model cards and API documentation for authoritative specifications, licensing and pricing.
