explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR — Pro, Fast, and whether the shared space is real
  • What people are asking
  • Official snapshot (from Cohere's table)
  • Other vendor scores worth logging, not tattooing
  • What to do this week
  • Honest limitations
  • Related on explainx.ai
← Back to blog

explainx / blog

Cohere Embed 5 Pro vs Fast: Shared Space, ViDoRe V3

Cohere, Embeddings, RAG, Vector Search, Enterprise AI

Cohere Embed 5 Pro is $0.12 and Fast $0.08 per million text tokens, sharing one index. Official ViDoRe V3: 85.8 vs 84.5 — still vendor scores.

Oct 1, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
Cohere Embed 5 Pro vs Fast: Shared Space, ViDoRe V3

Cohere published Introducing Embed 5 on September 30, 2026. The useful sentence is not "frontier." It is one index, two models: Embed 5 Pro and Fast share an embedding space so you can encode the corpus with the expensive model and hit it with the cheap one at query time.

List prices on that post, verified against the vendor page rather than a recap: $0.12 / 1M text tokens for Pro, $0.08 / 1M for Fast, $0.40 / 1M for images on both. Aggregators that repeated those figures were repeating Cohere.

The same day Perplexity previewed pplx-embed-v2-context-9b as a Hugging Face late-chunk text model with API still forthcoming. Different product: hosted multimodal SLA versus an open 9B contextual checkpoint. If you are choosing what to pay this week, start here. If you are choosing how to encode long prose with surrounding context, read the Perplexity piece too.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR — Pro, Fast, and whether the shared space is real

table · 2 cols
QuestionDirect answer
What shipped?Embed 5 Pro + Embed 5 Fast, GA on Cohere API / Model Vault / Foundry / SageMaker (Sep 30, 2026)
Official text price$0.12 Pro, $0.08 Fast per million tokens
Official image price$0.40 per million tokens, both tiers
Context / languages128K tokens; 100+ languages; text, images, fused text+image
Dims / formats2048, 1536, 1024, 768, 512, 256; float, int8, binary; Matryoshka
Shared space?Cohere says yes if output dimension matches (also with Matryoshka truncation and int8)
Recommended patternIndex Pro, query Fast
ViDoRe V3 (vendor, parsed text, RCP-nDCG@10)Pro 85.8, Fast 84.5 (Cohere: +8.8 vs Embed 4 for Pro)
Trust the leaderboard?No. Run your own first-stage Recall/nDCG. RCP-nDCG is a two-stage reorder

What people are asking

"Is $0.12 / $0.08 actually official?"

Yes, on cohere.com/blog/embed-5: "Pricing is $0.12 per million tokens for Pro and $0.08 per million tokens for Fast," with the snapshot table repeating those text rates and $0.40 image rates. Tao.media and other recaps matched that. Still confirm the pricing page when you sign a contract; blog tables drift.

"If they share a space, why buy Pro at all?"

Because mixed-model retrieval is not identical to Pro-Pro. Cohere's own 40-dataset matrix, normalized to Pro corpus + Pro query = 100:

table · 3 cols
Mean retrieval qualityCorpus: FastCorpus: Pro
Query: Fast96.698.4
Query: Pro97.3100

Index Fast + query Fast is the cheap rectangle (96.6). Index Pro + query Fast is the recommended rectangle (98.4). Index Fast + query Pro (97.3) is the awkward one: you paid latency on the query path and still indexed with the small model.

Cohere says no dataset in that suite showed a "major failure." That suite is theirs. Your clause-lookup corpus can still break if Pro and Fast disagree on a rare language or a table-heavy PDF.

Both sides must use the same output dimension. Mixing 1024-d Pro docs with 768-d Fast queries is not the product.

"Is ViDoRe V3 a clean win?"

ViDoRe V3 is ILLUIN Technology's enterprise visual document retrieval suite (with NVIDIA contribution): ~26,000 pages, ~3,099 queries, six languages, heavy human annotation. It is a real benchmark. Cohere's headline 85.8 / 84.5 is not "we submitted to MTEB and screenshotted the public board."

Cohere evaluated parsed text outputs curated by the ViDoRe authors, scored with RCP-nDCG@10. Footnote 1 on the launch post:

RCP-nDCG@10 requires evaluating embedding models in a two-stage retrieval setup, using their similarity scores to reorder a fixed candidate set. Scores therefore reflect reranking quality rather than first-stage retrieval performance.

So the viral table is closer to rerank-the-shortlist than search-the-whole-index. Their comparison set in that plot: Voyage 4 Large 83.7, Gemini Embedding 2 83.2, Embed 4 77.0, OpenAI text-embedding-3-large 75.5. Fast at 84.5 is the interesting claim: the cheap tier still sits above the named large competitors on this protocol.

Pro "leads five of eight domains" and ties Voyage 4 Large on energy, with largest Embed 4 gaps on HR (+11.4) and industrial (+10.3), still per Cohere.

If your production retriever is ANN over millions of chunks, demand Recall@k / nDCG@k on the full corpus, not only RCP-nDCG on a fixed candidate list.

"How is this different from Perplexity's contextual 9B?"

table · 3 cols
AxisCohere Embed 5Perplexity pplx-embed-v2-context-9b-preview
DateSep 30, 2026 GA APISep 30, 2026 HF preview
ModalityText, page images, fusedText late-chunk contextual
Context128KEval writes 32K single pass
BuyingHosted + vLLM privateWeights first, API later
TrickShared Pro/Fast spaceCompressor teacher vs gold chunk
Bench to quote carefullyViDoRe V3 RCP-nDCGPrivate context-bench + ConTEB

You can use both ideas in one stack: Cohere for multimodal enterprise PDFs, Perplexity-style late chunking for long text where pronouns and headers live far from the answer. Neither replaces agentic search on code.

Official snapshot (from Cohere's table)

table · 3 cols
CapabilityEmbed 5 ProEmbed 5 Fast
Best for (vendor)Max quality, offline indexing, complex corporaInteractive search, high-volume RAG, agents
Context128K128K
InputsText, images, fusedSame
Languages100+100+
Output dims2048 … 256Same
Formatsfloat, int8, binarySame
Self-hostYes (vLLM in the post)Yes
Text price$0.12 / 1M$0.08 / 1M
Image price$0.40 / 1M$0.40 / 1M

Model ID in their Python snippet: embed-v5.0-pro with input_type="search_document" / "search_query" and output_dimension=1024. Use the Fast ID from current docs when you wire the query path — do not guess a typo into production.

Other vendor scores worth logging, not tattooing

Finance (Cohere, RCP-nDCG@10 unless they labeled otherwise): FinanceBench 80.1 Pro / 80.0 Fast; FinQA 90.0 / 88.8; ViDoRe V3 Finance 85.0 / 83.9. They say Pro averages 3.3 points above the next non-Cohere competitor they name (Gemini Embedding 2) across that finance bundle, and +21.4 vs text-embedding-3-large on FinanceBench.

Parsed-document suite (PDFs parsed with Gemini 1.5 Flash in their protocol): Pro 84.8, Voyage 4 Large 83.6, Fast 83.4, Gemini Embedding 2 80.8, Embed 4 78.6.

Fused text-image: Pro 82.3 vs Fast 81.2 vs Gemini Embedding 2 61.3 (five datasets). Page-image finance-ish average: Pro 77.0, Fast 73.2, Embed 4 71.1, Voyage Multimodal 3.5 70.1, Gemini Embedding 2 56.7.

European language composite: Pro 77, Voyage 4 Large 76, Gemini Embedding 2 73, about +7 vs Embed 4. They also table ten more languages where Gemini or Voyage still win several rows (Japanese, Korean, Arabic, Telugu, Thai, and others in their grid). Do not sell Embed 5 as uniformly best multilingual; their own table contradicts that.

Throughput: Fast averages 2.4× document throughput vs Pro across ~200-token and ~1K-token contexts, per Cohere. That is indexer math, not query p50.

Storage: they repeat the Matryoshka arithmetic — 2048-d float32 = 8 KB, 1024-d int8 = 1 KB, 256-d binary = 32 bytes. Across 100M chunks they quote 819 GB → 3.2 GB of raw vectors. Graph indexes are extra. They recommend 1024-d int8 as the default efficiency point; binary as first-pass before a higher-precision rerank.

All of the above is Cohere-measured. Put it next to Q2D-Web if your traffic is agent-rewritten web queries, and next to a private 200-query gold set as the embeddings guide already tells you to build.

What to do this week

  1. Create a key and embed a slice with Pro documents and Fast queries at 1024 int8, same dimension both sides.
  2. Keep a Pro-Pro control on the same slice. If Fast queries drop more than Cohere's 1.6% relative mean on your nDCG@10, stop quoting 98.4.
  3. Separate text and image bills. $0.40/M image tokens will dominate if you embed page screenshots at 128K habitually.
  4. Do not retire the reranker on RCP-nDCG faith. If you already use Cohere Rerank or a cross-encoder, keep it until first-stage metrics move.
  5. Private deploy: the post says vLLM for VPC/on-prem. That is ops, not a quality fairy. Match input_type and dims to the API or your spaces diverge.
  6. Compare apples: Perplexity's contextual 9B is not a drop-in for page images. Cohere Parse (already in the dictionary as Cohere Parse) is the PDF-to-markdown sibling if you stay in that stack.
python
import os
import cohere
import numpy as np

co = cohere.ClientV2(api_key=os.environ["CO_API_KEY"])

docs = [
    "Net interest margin narrowed 12 bps to 2.61% as deposit costs rose.",
    "Torque the mounting bolts to 45 Nm in a star pattern before refitting the cover.",
]

doc_emb = co.embed(
    model="embed-v5.0-pro",
    input_type="search_document",
    texts=docs,
    output_dimension=1024,
    embedding_types=["float"],
).embeddings.float_

query_emb = co.embed(
    model="embed-v5.0-pro",
    input_type="search_query",
    texts=["What happened to net interest margin last quarter?"],
    output_dimension=1024,
    embedding_types=["float"],
).embeddings.float_[0]

# Swap the query model to Fast once you confirm the GA Fast model ID
# and keep output_dimension=1024.

D = np.array(doc_emb)
q = np.array(query_emb)
scores = D @ q / (np.linalg.norm(D, axis=1) * np.linalg.norm(q))
print(docs[int(np.argmax(scores))])

The snippet is adapted from Cohere's launch post (Pro on both sides). The production move they want is Pro documents, Fast queries, same output_dimension.

Honest limitations

  • Vendor benches, including ViDoRe V3 as they ran it. Parsed-text + RCP-nDCG is not page-image first-stage search.
  • Shared space is approximate. 1.6–3.4% mean loss in their matrix; tails unpublished.
  • Image tokens cost 5× Fast text. Multimodal RAG bills can invert the "Fast is cheaper" story.
  • 128K context is a capability, not a default. Encoding 100-page PDFs as one vector is still a pooling bet; chunking remains the embeddings guide default.
  • Integrations list (LangChain, Weaviate, Qdrant, and so on) is marketing adjacency. Your index still needs compatible dims and metric.
  • Compass Cloud private beta is a separate retrieval product. Do not assume Embed 5 GA includes Compass.

For a stale but structured closed-vs-open menu, see top 10 embedding models. Refresh it: Embed 4 is no longer the Cohere row to quote.

Related on explainx.ai

  • Perplexity pplx-embed-v2-context-9b-preview
  • What are embeddings? Vector search complete guide
  • Q2D-Web agentic RAG benchmark
  • RAG vs agentic RAG
  • Top 10 open and closed embedding models
  • Perplexity Fast Embeddings GPU serving
  • What is an embedding? Examples
  • Semantic vs vector vs hybrid search

Official sources: Cohere Embed 5 launch, ViDoRe V3 introduction, Cohere Embed product page.

Prices, model IDs, and benchmark figures are as of October 1, 2026, from Cohere's September 30, 2026 Embed 5 post. Re-check Cohere docs before you lock an index or a contract.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 1, 2026

Perplexity pplx-embed-v2-context-9b: Contextual Chunk Embeddings

On September 30, 2026 Perplexity Research and turbopuffer published Contextual embedding beyond the gold passage and a Hugging Face preview of pplx-embed-v2-context-9b-preview. The model encodes a document in one pass so each chunk vector sees surrounding context. API access is still coming. The reported ConTEB and private context-bench numbers are vendor benches.

Sep 7, 2026

Universal Geometry of Embeddings: Why "Safe" Vector Databases Aren’t

Researchers Jha, Zhang, Shmatikov, and Morris introduced an unsupervised method — now widely called vec2vec — for translating embeddings from one vector space into another without paired data or access to the original encoder. The security implication: leaked embedding vectors, long assumed to be effectively anonymized, can be translated into a known space and used to infer sensitive information about the underlying text.

Aug 27, 2026

Cohere Parse 5: Near-Frontier Document Parsing at $1.50/1k Pages

Cohere announced Parse 5 on August 27, 2026: a 2.3B vision parser that scores 79.2 on ParseBench (tables, faithfulness, semantic formatting) at $1.50 per 1,000 pages. It sits just under GPT-5.5 / Opus 4.8 / Gemini 3.5 Flash and well above hyperscaler OCR — with a free Hugging Face demo and Model Vault for high-volume work.