explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR — should you touch this preview?
  • What people are asking after the headline
  • How the training recipe actually works
  • What context-bench actually measures
  • ConTEB, domain suites, storage — still vendor benches
  • How this sits next to agentic RAG and Cohere Embed 5
  • What to do this week (if you actually retrieve documents)
  • Honest limitations
  • Related on explainx.ai
← Back to blog

explainx / blog

Perplexity pplx-embed-v2-context-9b: Contextual Chunk Embeddings

Perplexity, Embeddings, RAG, Vector Search, Retrieval

Perplexity's Sep 30 contextual embedder trains chunks from a compressor teacher. Hugging Face preview; API later. Treat the benches as vendor-run.

Oct 1, 2026·11 min read·Yash Thakker
add explainx.ai
go deep
Perplexity pplx-embed-v2-context-9b: Contextual Chunk Embeddings

On September 30, 2026, Perplexity Research and turbopuffer published Contextual embedding beyond the gold passage. The artifact is pplx-embed-v2-context-9b-preview: a 9B contextual embedding model that produces one vector per chunk after encoding the whole document.

If you already read explainx.ai's embeddings and vector-search guide, the pain is familiar. You split a 10-K, a lease, or a clinical protocol into windows. The sentence that answers the query no longer names the company, the revision, or the table header. Independent chunk embeddings then collide with near-duplicate sentences from the wrong sibling document.

The preview tries to train that away. The API is not the story this week. Hugging Face is. Vendor charts are still vendor charts.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR — should you touch this preview?

table · 2 cols
QuestionDirect answer
What shipped?Research post + HF preview of pplx-embed-v2-context-9b-preview (Sep 30, 2026)
What is new vs v1?Teacher is Perplexity's context compression model, not a single gold-chunk LLM label
Inference shapeOne document pass, learned chunk-separator token, mean-pool tokens per chunk
Dims / dtypeMatryoshka 2048 and 1024; native int8 via QAT; released as a checkpoint soup
Where to run itHugging Face preview. Perplexity says it is working on the API
Benches namedPrivate turbopuffer context-bench; public ConTEB; internal Q2C/Q2D suites
Headline vendor numbersAt K=10 on context-bench: 45.5% Answer, 40.6% Evidence, 31.1% All-Evidence; Document@1 15.2%, Document@10 61.6%
Should you rebuild production?Only after a private eval. Do not treat ConTEB average as a buying decision

What people are asking after the headline

"Is this just another pplx-embed checkpoint?"

No. Perplexity already ships independent pplx-embed-v1 models and contextual pplx-embed-context-v1 models in 0.6B and 4B sizes. Those are the ones in the contextualized embeddings API docs (pplx-embed-context-v1-0.6b at $0.008 / 1M tokens and pplx-embed-context-v1-4b at $0.05 / 1M, 32K context). This v2 preview is a 9B ColBERT-initialized student with a different training objective.

The serving story you already have on explainx.ai is Fast Embeddings on GPUs: Ivy, Tulip, ROSE, CUDA graphs. That post is how Perplexity runs embedders. This post is how they train a contextual one. Do not collapse the two.

"What is a gold passage and why would I care?"

Classic dense retrieval labels one chunk as relevant. Contrastive training then treats every other chunk in the same document as a negative. That is convenient for MS MARCO-style datasets. It is a bad match for documents where the answer sentence is true only after you also retrieve the definition, the revision banner, or the table header.

Perplexity's own sentence from the post:

A gold passage is often not enough in practice. It may contain the answer but lack the supporting context needed to understand or verify it.

That is the whole product pitch. If your RAG reader is an agent that should cite and verify, you want Answer@K and Evidence@K, not a single gold hit.

"Did they just invent late chunking?"

No. Late chunking — encode the document once, pool afterwards — is the standard contextual-embedding pattern. Perplexity says prior work (ConTEB and their own pplx-embed-context-v1) already showed gains over independent chunks on long documents. The claimed novelty is how labels are generated: distill token-level scores from a query-aware context compressor, then aggregate those scores onto whatever chunk boundaries you sampled that batch.

That is a training-time teacher. At inference there is no extra compressor and no extra reranker in the recipe they describe. Storage is still one vector per chunk.

"Can I call this from the Perplexity API this week?"

Perplexity writes, in the release section:

A preview of our model is available on Hugging Face, and all results in this post are reported for this preview. We are working on making the model available through the Perplexity API.

Until that lands, production callers stay on v1 IDs. Treat the 9B checkpoint as an eval toy unless you already self-host large embedders.

How the training recipe actually works

The student is an in-house 9B ColBERT retriever plus a linear projection to 2048 dimensions. Chunk boundaries are marked with a learned <|chunk_sep|> token. Chunk vectors are mean-pooled token embeddings. Queries go through the same model and are also mean-pooled.

Two losses share a batch:

  1. Document-level InfoNCE. Document score is the max chunk–query cosine (ColBERT's MaxSim idea, applied to chunks). In-batch documents are negatives.
  2. Chunk-level distillation. The compressor scores every token in the positive document. A chunk's teacher score is the mean of its top-n token scores. Softmax those scores into a target; other documents' chunks get zero. The student matches that distribution with forward KL.

Combined objective: weighted sum of the two. Batches are sampled from a single dataset so in-batch negatives are hard. Chunking strategy is resampled per batch, which is the point of token-level teachers: you do not re-annotate when you change window size.

Training data: roughly 430 public and in-house query–document datasets, 50+ languages. In-house pairs come from PII-filtered production data and synthetic queries over long documents. Perplexity says no ConTEB training data and no context-bench data during development. The released weights are a model soup of checkpoints from the same run.

That last sentence matters more than the marketing average. Soup + private bench + self-reported ConTEB is a lot of degrees of freedom. Reproduce on your leases before you throw out BM25.

What context-bench actually measures

turbopuffer holds context-bench private: 2,099 queries, 38,894 documents, sentence chunks totaling 2,458,072 vectors, 21 domains. Median primary target document is about 6,100 tokens; 1,061 of 1,197 distinct primary targets have at least 200 sentences. Eval is exhaustive over all chunks, so index config is not the confounder.

Three metrics, not one:

table · 2 cols
MetricWhat it asks
Document@KDid the right long document beat near-duplicate siblings?
Answer@KDid an accepted answer sentence land in the top K chunks corpus-wide?
Evidence Recall@KAfter dropping the top answer chunk, and only if the gold doc is in the top 10, how many evidence groups are recovered?

The lease example in the paper is the useful one. Many files share the sentence "Monthly rent is $X." The model has to use address and year context that may sit thousands of tokens away. That is closer to enterprise RAG than a Wikipedia gold passage.

Perplexity reports the preview leading Answer@K and Evidence Recall@K at every cutoff they plotted, plus All-Evidence@10. At K = 10: 45.5% answer recall, 40.6% evidence recall, 31.1% all-evidence. Document@1 15.2%, Document@10 61.6%. Versus voyage-context-4 at K=10: +14.4 points answer, +5.0 evidence.

Those are Perplexity's numbers on a bench turbopuffer holds. Other labs can submit to contextbench@turbopuffer.com. You cannot download the labels and rerun tonight. That is the contamination argument, and it is also why you should not paste 45.5% into a board deck as independent science.

ConTEB, domain suites, storage — still vendor benches

On ConTEB, Perplexity says the preview has the highest average nDCG@10 among the models in their figure. It does not win every task: they say pplx-embed-context-v1-4B is higher on NarrativeQA and Nemotron-3-8B is highest on COVID-QA (where lexical matching on medical terms can beat context). Average-without-per-task is how vendor plots get shared. Read the task bars.

On their internal query-to-chunk (Q2C) and query-to-document (Q2D) suites built from public datasets (LegalBench, FinanceBench, plus extra tasks), they say the preview leads average chunk retrieval and some domains, while voyage-context-4 is higher in finance and multilingual, and Nemotron-3-8B leads conversation. On document retrieval the preview is slightly below voyage-context-4 on average. If you only needed document IDs, a contextual 9B is not an automatic upgrade.

Storage claim, still theirs: contextualization does not multiply vector count versus the same chunker. Truncating 2048 → 1024 and using int8 drops bytes per vector. They say 1024-d int8 is 1 KB/vector and slightly beats voyage-context-4 at 2048-d float32 (8 KB/vector) on their chunk-retrieval plot. 2048-d int8 is 2 KB. Those figures are vector payload only, not HNSW graphs or metadata.

Chunk-size sweep: mean nDCG@10 across 74 MTEB tasks falls from 81.0% at 64-token chunks to 79.9% at 512 tokens. Modest. Useful if your indexer already picked a window for other reasons.

All of those evals encode documents up to 32,768 tokens in one pass. If your PDFs are longer, you still have a packing problem. The compressor teacher does not magically appear at query time.

How this sits next to agentic RAG and Cohere Embed 5

explainx.ai's RAG vs agentic RAG argument still holds for code: grep and structure beat naive chunk vectors. Contextual embeddings are an attempt to make chunk vectors less naive on long prose, filings, and manuals — the corpora where agents still retrieve then read.

Q2D-Web is a different axis: web-scale first-stage retrieval under agent-rewritten queries. A model that wins private long-doc evidence recall can still lose Recall@1000 on 190M web docs. Do not pick an embedder from one plot.

Same week, Cohere shipped Embed 5 Pro and Fast with a shared embedding space, 128K context, and vendor ViDoRe V3 scores. That is a hosted multimodal API with list prices. Perplexity's v2 is an open preview of a late-chunk text model with API forthcoming. Different buying motion. If you need images and a SLA this week, Cohere is the product. If you need to inspect a 9B contextual checkpoint, Perplexity is the lab drop.

For a closed-vs-open shortlist that is already stale on dates, see the top 10 embedding models roundup and refresh it with your own gold set. HyDE and Sentence Transformers v6 ColBERT remain the technique companions: hypothetical queries and late interaction are not replaced by one 9B soup.

What to do this week (if you actually retrieve documents)

  1. Do not wait for the API if you only wanted a paper. Read the post, note the teacher, note the private bench.
  2. If you self-host, pull the Hugging Face preview, encode a 100–200 query slice of your long docs, and score Answer plus a cheap evidence checklist (did the retrieved window contain the entity name?). Compare against independent chunks from the same backbone class and against v1 contextual if you already pay Perplexity.
  3. Match query encoding. Contextual doc models still need queries encoded the way the card says — v1 docs wrap queries as a one-element inner list. Confirm the v2 card before you mix spaces.
  4. Keep hybrid search. Contextual vectors do not retire BM25 on SKUs, error codes, or clause IDs. See the semantic vs hybrid search guide.
  5. Budget GPU like an LLM prefill. A 9B bidirectional pass over 32K tokens is not a MiniLM call. Perplexity's own GPU serving writeup is the honest infra companion.
  6. Log agent queries separately. If your users are agents, add a Q2D-Web-style split rather than only human FAQ queries.

Example shape for a contextualized v1 API call (still the documented production path, not the v2 preview):

bash
curl -X POST https://api.perplexity.ai/v1/contextualizedembeddings \
  -H "Authorization: Bearer $PPLX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "pplx-embed-context-v1-4b",
    "input": [[
      "Northlake is the registrant in this 10-K.",
      "The company repurchased $12.8 billion of shares."
    ]]
  }'

The v2 preview will not accept that ID. When the API ID lands, re-read the model card for encoding_format, Matryoshka dimensions, and whether queries share the contextual endpoint.

Honest limitations

  • Preview tag. Weights can move. Numbers in the blog are for this soup.
  • Private primary bench. You cannot audit label quality or contamination from the outside.
  • Teacher is Perplexity's compressor. Distillation quality is capped by that model. Errors in token relevance become chunk supervision.
  • 32K single pass. Longer docs still need packing, hierarchical retrieval, or an agent that greps.
  • Not multimodal. Page images, slides, and scanned tables are Cohere Embed 5 / ColPali territory, not this checkpoint.
  • ColBERT init, single-vector out. You get chunk vectors, not token-level late interaction at query time. Storage looks like a bi-encoder index.
  • Vendor domain plots. Finance and conversation losses versus named baselines are in their own figures. Copy those caveats into your eval notes.

Related on explainx.ai

  • Cohere Embed 5 Pro vs Fast and ViDoRe V3
  • Perplexity Fast Embeddings on GPUs — Ivy, Tulip, ROSE
  • Q2D-Web: agentic RAG retrieval benchmark
  • What are embeddings? Vector search complete guide
  • RAG vs agentic RAG
  • Top 10 open and closed embedding models
  • What is an embedding? Examples
  • Sentence Transformers v6 ColBERT

Official sources: Contextual embedding beyond the gold passage, Perplexity contextualized embeddings docs, pplx-embed Hugging Face collection.

Model IDs, benches, and API status are as of October 1, 2026, from Perplexity's September 30, 2026 Hub post. Preview weights and vendor scores can change; re-check Hugging Face and the Hub before you freeze an index.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 1, 2026

Cohere Embed 5 Pro vs Fast: Shared Space, ViDoRe V3

On September 30, 2026 Cohere launched Embed 5 in Pro and Fast tiers that share an embedding space, so you can index with Pro and query with Fast. List price is $0.12 vs $0.08 per million text tokens ($0.40 for images). ViDoRe V3 averages 85.8 and 84.5 are Cohere's own RCP-nDCG@10 numbers.

Sep 10, 2026

Perplexity Q2D-Web: 190M-Doc Benchmark for Agentic RAG Retrieval

On September 10, 2026, Perplexity published Q2D-Web (Query2Doc-Web) — a large-scale benchmark for first-stage retrieval in agentic RAG systems. Built from 23,000 PII-free production searches over nine months, it pairs 190 million web documents with 69,721 agent-reformulated queries, ten languages, and an average of 99.6 positive relevance judgments per query. explainx.ai breaks down why the benchmark exists, how it differs from MS MARCO Web, and what it means if you ship embedding models or agent search stacks.

Sep 7, 2026

Universal Geometry of Embeddings: Why "Safe" Vector Databases Aren’t

Researchers Jha, Zhang, Shmatikov, and Morris introduced an unsupervised method — now widely called vec2vec — for translating embeddings from one vector space into another without paired data or access to the original encoder. The security implication: leaked embedding vectors, long assumed to be effectively anonymized, can be translated into a known space and used to infer sensitive information about the underlying text.