explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • How the filter sits in the pipeline
  • Does Jev replace RAG?
  • What people are asking: the 73% headline
  • Cost and latency: Jev vs BM25 vs embeddings
  • How to set CONTEXT_FILTER
  • When to keep embeddings
  • Python 3.12+ and the rest of v3.7.0
  • Reproduce the 28-task replay
  • Honest limitations
  • What this means for builders
  • Related on explainx.ai
← Back to blog

explainx / blog

GPT Researcher 3.7.0: Jev Scores Passages by Usefulness

GPT Researcher, Jev, TypeSafe AI, RAG, AI Agents

GPT Researcher v3.7.0 scores scraped passages with Jev. Default without TYPESAFE_API_KEY is BM25. The 73% figure is kept-chunk precision on 28 tasks.

Sep 29, 2026·10 min read·Yash Thakker
add explainx.ai
go deep
GPT Researcher 3.7.0: Jev Scores Passages by Usefulness

Assaf Elovic's GPT Researcher tagged v3.7.0 on September 26, 2026 as "Jev update and many bug fixes." The headline feature is a context filter: after scrape, each chunk is scored by Jev (TypeSafe AI's System One model) on usefulness for the sub-query, not cosine similarity to the query embedding. The v3.7.0 release notes land feat #2161 and feat #2162. Official docs live at docs.gptr.dev/docs/gpt-researcher/gptr/context-filter (verified live: title "Context Filter | GPT Researcher").

If you saw an aggregator line like "73% relevant context," treat it as a compressed precision number, not a product-quality score. Primary eval notes say 73% of chunks Jev kept were judged relevant versus 46% for embeddings on 28 replayed tasks. Users without a TypeSafe key do not get that Jev precision; they get BM25.

explainx.ai already covers TypeSafe's Jev launch, wiring Jev into agent routing, and cheap verification checkpoints. This post is the research-agent instance of that pattern: score retrieved passages before the writer spends tokens on them.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 2 cols
QuestionAnswer
What shipped?GPT Researcher v3.7.0 (PyPI gpt-researcher 0.16.0 in the release notes)
Does Jev replace RAG?No. Filter after scrape. Vector stores and CONTEXT_FILTER=embeddings still exist
Default with no TypeSafe key?BM25 keyword, not embeddings (#2162)
What is 73%?Kept-passage precision for Jev vs 46% embeddings on 28 tasks, not report accuracy
Keyword vs embeddings?BM25 matched or beat embeddings on the published measures; directional 14–6 head-to-head
Python?3.12+ required
How to opt into old path?CONTEXT_FILTER=embeddings
Where is the eval?evals/context_filter/

How the filter sits in the pipeline

Search still finds URLs. Scrapers still pull pages. The new step is select_context() in gpt_researcher/context/select.py: each page is capped at 50,000 characters, split into 1,000-character chunks with 100 characters of overlap, then ranked.

Jev and embeddings keep up to 10 chunks per sub-query. Keyword keeps up to 25. Sub-queries run concurrently; joined passages go to the writer. Report chat uses the same function, treating the finished report as one page and the latest message as the query.

Small page sets (under 8,000 characters total, and no more pages than the chunk budget) skip filtering in every mode.

Retrieval and context selection as a pipeline, reused as the hero for GPT Researcher's Jev usefulness filter versus embedding similarity

That is closer to a reranker than to "RAG is dead." If you already think in RAG versus MCP retrieval terms: GPT Researcher still retrieves; it changed how it prunes.

Does Jev replace RAG?

No. Three facts from the docs and PRs:

  1. Jev never generates the report. It returns a 0–3 usefulness score (probability-weighted rubric). The writer LLM still writes.
  2. Embeddings are optional, not deleted. CONTEXT_FILTER=embeddings restores cosine similarity. A caller-supplied vector_store= still uses embeddings. Detailed-report section overlap still uses embeddings when a model is available, otherwise keywords.
  3. Without TYPESAFE_API_KEY, you are not on Jev. auto falls through to BM25. That is the opposite of "everyone now runs a System One RAG stack."

The useful comparison is Jev as a typed usefulness judge versus embedding similarity as a topical neighbor search. TypeSafe's own product story is the same split explainx.ai documented at launch: Jev does not replace chat or long-form generation. GPT Researcher used that primitive on chunks, the same family of decision as agent routing and post-retrieval checkpoints.

What people are asking: the 73% headline

Official copy: Jev's kept context is "59% more relevant than embeddings, at the same cost (73% of kept passages relevant, against 46%)." Relative lift: (73 − 46) / 46 ≈ 59%. Absolute: 73 vs 46, not "the product is 73% accurate."

table · 2 cols
Claim you might seeWhat the eval actually says
"73% relevant context"Precision of kept chunks, judged by gpt-5.4-mini
Implies all usersOnly Jev mode with a working TypeSafe key
Implies BM25 is 73%Keyword shipped at 51% precision (relative-threshold, up to 25)
Implies better answersSimpleQA 18–20 / 20 for every filter — saturated
Implies huge eval28 tasks: 20 SimpleQA + 8 open-ended, one writer (gpt-5.4), one report type

Head-to-head (blind, both orders, win only if both agree): Jev 15 · 10 · 3 vs embeddings (sign test on 18 non-ties, p ≈ 0.008 in the write-up). Keyword 14 · 8 · 6 (p ≈ 0.12 — directional, not as strong).

Unranked control (10 chunks round-robin, no ranking): 33% precision. Forced Jev top-10 with no threshold: 50% — the 1.5 min score is where most of the Jev gain comes from. Cosine similarity cannot say "nothing else on this page is worth including."

Do not treat 73% as a substitute for reading evals/context_filter/README.md. Caveats in that README: 28 questions, same model family writes and judges, detailed/deep research not replayed.

Cost and latency: Jev vs BM25 vs embeddings

Numbers below are from the shipped docs table and README footnotes (median over the 28-task replay unless noted). Writer time was about 45 seconds for every filter; total run time moves mainly with the filter step. Docs put a typical research run around two minutes, with Jev adding about 0.7s at the median versus embeddings.

table · 5 cols
ModeKept-passage precisionFilter time (median)Context sent (median)Cost per report
Jev73%1.7s (p90 3.6s at concurrency 64)4.5k tokens$0.115
Keyword (BM25)51%0.02s6.4k tokens$0.116
Embeddings46%1.0s (p90 1.3s)7.6k tokens$0.117
Unranked control33%~0.01s8.0k tokens$0.117
No filtern/a—33.4k tokens (max 108k)$0.192

Jev API cost. Docs: Jev bills $0.042 per million input tokens, output free — the same published rate explainx.ai used in checkpoint cost math. Each chunk is its own POST https://api.typesafe.ai/v1/systemone call. Filtering added about half a cent per report in the benchmark; README tracked $0.59 of Jev across all configurations in the full eval spend (~$41 writing/filtering plus judging).

Throughput. Default JEV_CONCURRENCY=64 (shared per event loop). Warm call ~0.33s; ~0.45s at 13k tokens; 64 concurrent calls ~1.1s. Raising to 128 did not make filtering faster. HTTP 429/529 retry up to three times; other failures raise JevError and select_context falls back to keyword.

BM25. Pure Python in gpt_researcher/context/lexical.py, no extra dependency. ~20ms per typical sub-query. k1 = 1.5, b = 0.75. Keep chunks scoring at least 50% of the best (KEYWORD_RELATIVE_THRESHOLD), up to 25. Plain top-10 BM25 lost to embeddings (40% precision, 7–12 head-to-head). The relative threshold is the shipped default for a reason.

No filter. Wins open-ended pairwise vs embeddings (8–0 on the eight open-ended tasks) because more material reaches the writer, but costs ~65% more per report on average (~83% on open-ended) and is not faster. Fact-finding (SimpleQA) does not benefit.

For TypeSafe's broader speed/cost marketing, keep explainx.ai's Jev claims fact-check in mind: GPT Researcher's table is a narrow filter benchmark, not a 200× LLM replacement study.

How to set CONTEXT_FILTER

bash
export TYPESAFE_API_KEY=...              # enables jev under default auto
export CONTEXT_FILTER=auto               # auto | jev | keyword | embeddings | none

export JEV_MIN_SCORE=1.5                 # 0-3 usefulness a chunk needs
export JEV_CONCURRENCY=64
export JEV_CHUNK_SIZE=1000
export JEV_MODEL=jev-latest

export KEYWORD_RELATIVE_THRESHOLD=0.5
export KEYWORD_MAX_RESULTS=25

export SIMILARITY_THRESHOLD=0.42         # embeddings mode only
export COMPRESSION_THRESHOLD=8000        # small page sets pass through
table · 3 cols
ModeHow passages are chosenNeeds
auto (default)Jev if TYPESAFE_API_KEY is set, else keywordnothing
jevUsefulness score ≥ JEV_MIN_SCORE, up to 10TYPESAFE_API_KEY
keywordBM25 with relative threshold, up to 25nothing
embeddingsCosine similarity vs query embeddingan EMBEDDING provider
noneFull pagesnothing

Fallback chain: jev → keyword on missing key, network error, rate limit after retries, or malformed response. embeddings → keyword if the embedding model cannot be built. Only embeddings mode constructs an embedding model for standard research.

Jev's rubric (from docs): unrelated; same topic but does not help; partially answers or supporting facts; directly answers with specific facts. The response score is the probability-weighted level from 0 to 3.

http
POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer $TYPESAFE_API_KEY

{
  "state": "<the 1,000-character chunk>",
  "model": "jev-latest",
  "questions": {
    "usefulness": {
      "type": "score",
      "instructions": "How useful is this passage for answering the question: <sub-query>",
      "criteria": [
        "Unrelated to the question",
        "Same topic, but does not help answer the question",
        "Partially answers the question or gives useful supporting facts",
        "Directly answers the question with specific facts"
      ]
    }
  }
}

For schema-level wiring outside this repo, use the Jev integration guide and how Jev's Score primitive works.

When to keep embeddings

Keep CONTEXT_FILTER=embeddings when:

  • You already pay for an embedding endpoint and want no TypeSafe dependency.
  • Queries are paraphrases of source language (BM25's light stemming will miss a lot of synonymy).
  • You pass a custom vector_store — that path is still embedding-backed.
  • You are A/B testing against a corpus where cosine already works and you do not want a 28-task web-research replay to decide for you.

Prefer keyword when you want zero extra APIs, fastest filter (~50× vs embeddings on the published p50), and the maintainers' claim that BM25 matched or beat embeddings on every measure in this replay — remembering the pairwise result is weaker than Jev's.

Prefer Jev when you have a TypeSafe key, accept network + ~1.7s median filter time, and want the thresholded usefulness behavior. It is the same "cheap structured check after retrieval" idea as pipeline checkpoints, not a new retrieval index.

Prefer none only if you have budget for ~65%+ higher writer cost and your report type is open-ended enough that more raw pages help. Deep/detailed research multiply sub-queries; the docs warn the token gap grows there.

Python 3.12+ and the rest of v3.7.0

The GitHub release lists Python 3.12+ as a change to note. Pin your runtime before pip install -U gpt-researcher.

Also in the tag: retriever plugins via the gpt_researcher.retrievers entry point (#2154), frontend WebSocket double-connect fix (#2159), multi-agent section overlap (#2158), self-hosted frontend assets without third-party CDNs (#2082), plus 25+ community retriever/scraper/deep-research fixes.

#2161 originally defaulted auto without a TypeSafe key to embeddings. #2162 replaced that fallback with keyword so no embeddings provider is required for standard research. If you upgraded from a mid-PR mental model, re-read auto.

Reproduce the 28-task replay

bash
export OPENAI_API_KEY=... TAVILY_API_KEY=... TYPESAFE_API_KEY=...
python -m evals.context_filter.collect --out runs/
python -m evals.context_filter.replay  --runs runs/ --out results/
python -m evals.context_filter.judge   --results results/

Committed artifacts are scores and metrics in evals/context_filter/results/, not scraped pages or generated reports. Collect is documented at roughly $0.40 / 20 minutes in the README; the full configuration sweep was about $41 tracked plus $8–15 judging.

Honest limitations

  • n = 28, one writer, one report type. Not a leaderboard for every research agent.
  • LLM-as-judge precision and same-family pairwise judging. Self-preference is argued to apply equally, but it is still model-on-model.
  • Jev needs TypeSafe uptime. Failures silently become BM25. That is good for availability; it is bad if you assumed every run was Jev-scored.
  • Jev is not "never wrong." Schema-valid scores can still be bad relevance calls — the same distinction as launch coverage.
  • Keyword is not Jev. Do not quote 73% for a default Docker box with no TYPESAFE_API_KEY.

What this means for builders

If you run GPT Researcher as a research tool, v3.7.0 is a filter swap, not a RAG funeral. Set TYPESAFE_API_KEY if you want usefulness scoring; otherwise you already have a local BM25 default that the authors argue is good enough to drop mandatory embeddings. If you build agents, treat this as another production example of Jev as a typed gate after retrieval, next to routing and checkpoints — then measure on your corpus, because 28 web-research tasks will not match a private wiki.

Related on explainx.ai

  • TypeSafe AI launches Jev
  • How to wire Jev into agent routing
  • Jev as cheap verification checkpoints
  • How Jev works: RLCD and System One
  • Jev speed and cost claims, fact-checked
  • RAG vs MCP: complete comparison
  • Jev waitlist is gone

Official: v3.7.0 release · Context Filter docs · PR #2161 · PR #2162 · evals/context_filter · TypeSafe Jev intro


Version numbers, CONTEXT_FILTER defaults, Python 3.12+ requirement, and eval tables reflect GPT Researcher v3.7.0 and the Context Filter docs as of September 29, 2026. The 28-task replay is the authors' own; rerun evals/context_filter before treating 73% as a property of your workload.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 21, 2026

DocJev: LlamaIndex's Jerry Liu Puts Jev on Document Classification and Splitting

Jerry Liu, LlamaIndex's cofounder and CEO, released DocJev — an open-source library that hands document classification and document-splitting decisions to Jev instead of a general-purpose LLM. The published benchmark shows classification dropping from 794ms to 138.6ms median latency, but a replier's qualifier about the accuracy pilot's small sample size is worth reading before trusting it unattended.

Sep 21, 2026

Using Jev as Cheap Verification Checkpoints in Agent Pipelines

A checkpoint that costs a fraction of a cent only pays for itself if it changes what happens next. This guide works through where to place Jev checks in a research-to-article agent pipeline, the real cost math behind "cheap enough to check constantly," and the honest failure modes — noisy alarms, distracting context, and checks with no attached action — that make a checkpoint worthless even when it's nearly free.

Sep 21, 2026

Jev Ultrafast: Browser Use Puts Jev in the Browser Agent Loop

Browser Use, the team behind the popular browser-use agent library, shipped Jev Ultrafast — an open-source browser agent that reads a structured element table instead of screenshots and lets Jev pick an operation and a target element per step, with a small LLM only invoked to write text. The published demo completes a real Google Flights search in 7.1 seconds, with independent outcome verification.