explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

  1. Home
  2. /
  3. Dictionary
  4. /
  5. Q2D-Web
Evaluation & Benchmarksaka Query2Doc-Webaka Q2D Web

Q2D-Web

Q2D-Web is Perplexity's large-scale benchmark for first-stage retrievers in agentic RAG, pairing 190 million web documents with ~70,000 agent-reformulated queries and dense multi-source relevance labels.

Ask Melo about this← all terms

Released September 10, 2026 alongside arXiv:2609.08887, Q2D-Web (Query2Doc-Web) extends Perplexity's earlier Q2D embedding benchmark to web scale. Queries are reformulated from 23,000 PII-free production searches over nine months across ten languages (English ~65.8%). Three fixed judgment sets — agent citations, production web rankings, and a combined LLM-judged union — average about 99.6 positive labels per query in the combined view. A public Hugging Face leaderboard accepts open embedder submissions; the paper documents subcorpus sampling to cut full-corpus eval cost.

Related terms

AI BenchmarkRecallLLM-as-JudgeRed TeamingROC AUCFew-Shot Evaluation