Cohere published Introducing Embed 5 on September 30, 2026. The useful sentence is not "frontier." It is one index, two models: Embed 5 Pro and Fast share an embedding space so you can encode the corpus with the expensive model and hit it with the cheap one at query time.
List prices on that post, verified against the vendor page rather than a recap: $0.12 / 1M text tokens for Pro, $0.08 / 1M for Fast, $0.40 / 1M for images on both. Aggregators that repeated those figures were repeating Cohere.
The same day Perplexity previewed pplx-embed-v2-context-9b as a Hugging Face late-chunk text model with API still forthcoming. Different product: hosted multimodal SLA versus an open 9B contextual checkpoint. If you are choosing what to pay this week, start here. If you are choosing how to encode long prose with surrounding context, read the Perplexity piece too.
TL;DR — Pro, Fast, and whether the shared space is real
| Question | Direct answer |
|---|---|
| What shipped? | Embed 5 Pro + Embed 5 Fast, GA on Cohere API / Model Vault / Foundry / SageMaker (Sep 30, 2026) |
| Official text price | $0.12 Pro, $0.08 Fast per million tokens |
| Official image price | $0.40 per million tokens, both tiers |
| Context / languages | 128K tokens; 100+ languages; text, images, fused text+image |
| Dims / formats | 2048, 1536, 1024, 768, 512, 256; float, int8, binary; Matryoshka |
| Shared space? | Cohere says yes if output dimension matches (also with Matryoshka truncation and int8) |
| Recommended pattern | Index Pro, query Fast |
| ViDoRe V3 (vendor, parsed text, RCP-nDCG@10) | Pro 85.8, Fast 84.5 (Cohere: +8.8 vs Embed 4 for Pro) |
| Trust the leaderboard? | No. Run your own first-stage Recall/nDCG. RCP-nDCG is a two-stage reorder |
What people are asking
"Is $0.12 / $0.08 actually official?"
Yes, on cohere.com/blog/embed-5: "Pricing is $0.12 per million tokens for Pro and $0.08 per million tokens for Fast," with the snapshot table repeating those text rates and $0.40 image rates. Tao.media and other recaps matched that. Still confirm the pricing page when you sign a contract; blog tables drift.
"If they share a space, why buy Pro at all?"
Because mixed-model retrieval is not identical to Pro-Pro. Cohere's own 40-dataset matrix, normalized to Pro corpus + Pro query = 100:
| Mean retrieval quality | Corpus: Fast | Corpus: Pro |
|---|---|---|
| Query: Fast | 96.6 | 98.4 |
| Query: Pro | 97.3 | 100 |
Index Fast + query Fast is the cheap rectangle (96.6). Index Pro + query Fast is the recommended rectangle (98.4). Index Fast + query Pro (97.3) is the awkward one: you paid latency on the query path and still indexed with the small model.
Cohere says no dataset in that suite showed a "major failure." That suite is theirs. Your clause-lookup corpus can still break if Pro and Fast disagree on a rare language or a table-heavy PDF.
Both sides must use the same output dimension. Mixing 1024-d Pro docs with 768-d Fast queries is not the product.
"Is ViDoRe V3 a clean win?"
ViDoRe V3 is ILLUIN Technology's enterprise visual document retrieval suite (with NVIDIA contribution): ~26,000 pages, ~3,099 queries, six languages, heavy human annotation. It is a real benchmark. Cohere's headline 85.8 / 84.5 is not "we submitted to MTEB and screenshotted the public board."
Cohere evaluated parsed text outputs curated by the ViDoRe authors, scored with RCP-nDCG@10. Footnote 1 on the launch post:
RCP-nDCG@10 requires evaluating embedding models in a two-stage retrieval setup, using their similarity scores to reorder a fixed candidate set. Scores therefore reflect reranking quality rather than first-stage retrieval performance.
So the viral table is closer to rerank-the-shortlist than search-the-whole-index. Their comparison set in that plot: Voyage 4 Large 83.7, Gemini Embedding 2 83.2, Embed 4 77.0, OpenAI text-embedding-3-large 75.5. Fast at 84.5 is the interesting claim: the cheap tier still sits above the named large competitors on this protocol.
Pro "leads five of eight domains" and ties Voyage 4 Large on energy, with largest Embed 4 gaps on HR (+11.4) and industrial (+10.3), still per Cohere.
If your production retriever is ANN over millions of chunks, demand Recall@k / nDCG@k on the full corpus, not only RCP-nDCG on a fixed candidate list.
"How is this different from Perplexity's contextual 9B?"
| Axis | Cohere Embed 5 | Perplexity pplx-embed-v2-context-9b-preview |
|---|---|---|
| Date | Sep 30, 2026 GA API | Sep 30, 2026 HF preview |
| Modality | Text, page images, fused | Text late-chunk contextual |
| Context | 128K | Eval writes 32K single pass |
| Buying | Hosted + vLLM private | Weights first, API later |
| Trick | Shared Pro/Fast space | Compressor teacher vs gold chunk |
| Bench to quote carefully | ViDoRe V3 RCP-nDCG | Private context-bench + ConTEB |
You can use both ideas in one stack: Cohere for multimodal enterprise PDFs, Perplexity-style late chunking for long text where pronouns and headers live far from the answer. Neither replaces agentic search on code.
Official snapshot (from Cohere's table)
| Capability | Embed 5 Pro | Embed 5 Fast |
|---|---|---|
| Best for (vendor) | Max quality, offline indexing, complex corpora | Interactive search, high-volume RAG, agents |
| Context | 128K | 128K |
| Inputs | Text, images, fused | Same |
| Languages | 100+ | 100+ |
| Output dims | 2048 … 256 | Same |
| Formats | float, int8, binary | Same |
| Self-host | Yes (vLLM in the post) | Yes |
| Text price | $0.12 / 1M | $0.08 / 1M |
| Image price | $0.40 / 1M | $0.40 / 1M |
Model ID in their Python snippet: embed-v5.0-pro with input_type="search_document" / "search_query" and output_dimension=1024. Use the Fast ID from current docs when you wire the query path — do not guess a typo into production.
Other vendor scores worth logging, not tattooing
Finance (Cohere, RCP-nDCG@10 unless they labeled otherwise): FinanceBench 80.1 Pro / 80.0 Fast; FinQA 90.0 / 88.8; ViDoRe V3 Finance 85.0 / 83.9. They say Pro averages 3.3 points above the next non-Cohere competitor they name (Gemini Embedding 2) across that finance bundle, and +21.4 vs text-embedding-3-large on FinanceBench.
Parsed-document suite (PDFs parsed with Gemini 1.5 Flash in their protocol): Pro 84.8, Voyage 4 Large 83.6, Fast 83.4, Gemini Embedding 2 80.8, Embed 4 78.6.
Fused text-image: Pro 82.3 vs Fast 81.2 vs Gemini Embedding 2 61.3 (five datasets). Page-image finance-ish average: Pro 77.0, Fast 73.2, Embed 4 71.1, Voyage Multimodal 3.5 70.1, Gemini Embedding 2 56.7.
European language composite: Pro 77, Voyage 4 Large 76, Gemini Embedding 2 73, about +7 vs Embed 4. They also table ten more languages where Gemini or Voyage still win several rows (Japanese, Korean, Arabic, Telugu, Thai, and others in their grid). Do not sell Embed 5 as uniformly best multilingual; their own table contradicts that.
Throughput: Fast averages 2.4× document throughput vs Pro across ~200-token and ~1K-token contexts, per Cohere. That is indexer math, not query p50.
Storage: they repeat the Matryoshka arithmetic — 2048-d float32 = 8 KB, 1024-d int8 = 1 KB, 256-d binary = 32 bytes. Across 100M chunks they quote 819 GB → 3.2 GB of raw vectors. Graph indexes are extra. They recommend 1024-d int8 as the default efficiency point; binary as first-pass before a higher-precision rerank.
All of the above is Cohere-measured. Put it next to Q2D-Web if your traffic is agent-rewritten web queries, and next to a private 200-query gold set as the embeddings guide already tells you to build.
What to do this week
- Create a key and embed a slice with Pro documents and Fast queries at 1024 int8, same dimension both sides.
- Keep a Pro-Pro control on the same slice. If Fast queries drop more than Cohere's 1.6% relative mean on your nDCG@10, stop quoting 98.4.
- Separate text and image bills. $0.40/M image tokens will dominate if you embed page screenshots at 128K habitually.
- Do not retire the reranker on RCP-nDCG faith. If you already use Cohere Rerank or a cross-encoder, keep it until first-stage metrics move.
- Private deploy: the post says vLLM for VPC/on-prem. That is ops, not a quality fairy. Match
input_typeand dims to the API or your spaces diverge. - Compare apples: Perplexity's contextual 9B is not a drop-in for page images. Cohere Parse (already in the dictionary as Cohere Parse) is the PDF-to-markdown sibling if you stay in that stack.
import os
import cohere
import numpy as np
co = cohere.ClientV2(api_key=os.environ["CO_API_KEY"])
docs = [
"Net interest margin narrowed 12 bps to 2.61% as deposit costs rose.",
"Torque the mounting bolts to 45 Nm in a star pattern before refitting the cover.",
]
doc_emb = co.embed(
model="embed-v5.0-pro",
input_type="search_document",
texts=docs,
output_dimension=1024,
embedding_types=["float"],
).embeddings.float_
query_emb = co.embed(
model="embed-v5.0-pro",
input_type="search_query",
texts=["What happened to net interest margin last quarter?"],
output_dimension=1024,
embedding_types=["float"],
).embeddings.float_[0]
# Swap the query model to Fast once you confirm the GA Fast model ID
# and keep output_dimension=1024.
D = np.array(doc_emb)
q = np.array(query_emb)
scores = D @ q / (np.linalg.norm(D, axis=1) * np.linalg.norm(q))
print(docs[int(np.argmax(scores))])
The snippet is adapted from Cohere's launch post (Pro on both sides). The production move they want is Pro documents, Fast queries, same output_dimension.
Honest limitations
- Vendor benches, including ViDoRe V3 as they ran it. Parsed-text + RCP-nDCG is not page-image first-stage search.
- Shared space is approximate. 1.6–3.4% mean loss in their matrix; tails unpublished.
- Image tokens cost 5× Fast text. Multimodal RAG bills can invert the "Fast is cheaper" story.
- 128K context is a capability, not a default. Encoding 100-page PDFs as one vector is still a pooling bet; chunking remains the embeddings guide default.
- Integrations list (LangChain, Weaviate, Qdrant, and so on) is marketing adjacency. Your index still needs compatible dims and metric.
- Compass Cloud private beta is a separate retrieval product. Do not assume Embed 5 GA includes Compass.
For a stale but structured closed-vs-open menu, see top 10 embedding models. Refresh it: Embed 4 is no longer the Cohere row to quote.
Related on explainx.ai
- Perplexity pplx-embed-v2-context-9b-preview
- What are embeddings? Vector search complete guide
- Q2D-Web agentic RAG benchmark
- RAG vs agentic RAG
- Top 10 open and closed embedding models
- Perplexity Fast Embeddings GPU serving
- What is an embedding? Examples
- Semantic vs vector vs hybrid search
Official sources: Cohere Embed 5 launch, ViDoRe V3 introduction, Cohere Embed product page.
Prices, model IDs, and benchmark figures are as of October 1, 2026, from Cohere's September 30, 2026 Embed 5 post. Re-check Cohere docs before you lock an index or a contract.
