Cloudflare added a Web Search API to AI Gateway on October 2, 2026. One request returns structured search results from Ceramic.ai, Exa or Linkup, billed at the partners' list prices with no added markup and logged alongside your model calls.
For anyone building agents, the interesting part is not that search exists — every agent framework has search tools — but that it now shares a control plane with inference. One gateway, one bill, one log stream.
TL;DR — what people are asking
| Question | Answer |
|---|---|
| What shipped? | A Web Search API running through AI Gateway |
| Providers? | Ceramic.ai, Exa, Linkup |
| How do I call it? | REST endpoint or the Workers env.AI.websearch binding |
| Parameters? | query, provider, limit, plus the gateway ID |
| Cost? | Partner list API prices, no markup, drawn from AI Gateway credits |
| BYOK? | Yes, bring your own provider key |
| Logging? | Requests appear in normal AI Gateway logs and analytics |
| Privacy? | Zero Data Retention partners are identified; partners must follow bot-crawling standards |
| Status? | Cloudflare's post does not say beta; one tracker calls it open beta |
| Coming? | Native server tools in AI Gateway, with web search first |
What it is
Agents need live information. A model's knowledge stops at its training cutoff, so anything current — prices, docs, news — has to come from a tool. Cloudflare's Web Search API provides that tool through the same gateway you may already use for inference. It returns titles, URLs and descriptions, which you pass into the model's context.
The key design choice is putting search behind AI Gateway. That brings the controls developers already rely on: access control, analytics, billing and logs, applied to search as well as model calls. If you run agents on Cloudflare's platform, you can now keep inference, search and execution in one place.
How to call it
REST
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/websearch/ \
--request POST \
--header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
--header "Content-Type: application/json" \
--data '{
"query": "What are some fun things to do in Salt Lake City as fall approaches?",
"provider": "ceramic",
"limit": 5,
"options": {
"gateway": { "id": "default" }
}
}'
From a Worker
const response = await env.AI.websearch({
gatewayId: "default",
query: "What are some fun things to do in Salt Lake City as fall approaches?",
provider: "exa",
limit: 5,
})
const results = await response.json()
The parameters are the same in both: a query string, a provider, a result limit, and the gateway identifier.
The pieces that matter for production
| Feature | Why it matters |
|---|---|
| List-price billing, no markup | You pay what the provider charges; compare against calling the provider directly |
| Single credit balance | Search and inference draw from the same AI Gateway credits |
| Gateway logs and analytics | You can see search volume, errors and cost next to model calls |
| BYOK | Use your own provider contract and rate limits |
| ZDR identification | Zero Data Retention partners are marked, which helps with privacy reviews |
| Crawling standards | Partners must follow bot-crawling standards, which matters for the health of the web ecosystem |
If you already use AI Gateway to route across model providers, as many teams do — see how open models grew to most of Vercel AI Gateway's token volume for the broader gateway trend — then adding search to the same path reduces the number of vendors in your agent's critical path.
Which provider should you pick?
The announcement does not rank the three, and I have not benchmarked them. Choose by testing:
- Collect 30 real queries from your agent logs.
- Run each through all three providers with the same
limit. - Score relevance and freshness by hand, and note which returns usable descriptions.
- Compare latency and cost from the gateway analytics.
- Check the data-retention posture for your compliance needs, using the ZDR identification.
Because switching is a one-word change to the provider field, a provider bake-off costs little. Keep the choice in config, not code.
Where this fits in an agent
A typical flow:
- The agent decides it needs current information.
- It calls search with a focused query.
- Your code filters and ranks results, fetches the most relevant pages if needed, and passes snippets to the model.
- The model answers with citations.
Two cautions. First, search results are untrusted input. A web page can contain instructions aimed at your model, so treat retrieved text as data and apply the defenses in our guide to indirect prompt injection in AI agents. Second, limit what the agent can do after reading the web. An agent that can search and also execute code or spend money should run in a sandbox with scoped credentials; see Cloudflare Sandbox SDK 1.0 for one approach, and Cloudflare Wallets for spending controls.
Native server tools: the next step
Cloudflare says native server tools are coming to AI Gateway, starting with web search. Today you write the loop: call the model, see it ask for search, call the API, return results. A native tool would let the gateway handle that exchange. If it ships as described, it simplifies agent code and moves more behavior into the gateway, which raises a design question worth thinking about now: how much of your agent's logic do you want to live in a vendor's gateway versus your own code?
For the protocol side of tool calling, our MCP guide explains the standard approach, and what is an agent harness covers where tools sit in the stack.
What people are asking
Is this better than using Exa or Linkup directly?
It is the same providers behind a Cloudflare control plane. The advantage is unified billing, logging and access control. The trade-off is another layer between you and the provider. If you only use one provider and already have a contract, direct may be simpler.
Will it work with any model?
The search call is independent of the model. You call search, then pass the results to whichever model you use. The coming native tools would couple the two more tightly.
Does it work outside Workers?
Yes. The REST endpoint works from any backend with an account ID and API token.
What about rate limits and quotas?
I did not find limits in the sources I reviewed. Check the docs, and use BYOK if you need the provider's own limits.
Does it replace AI Search?
No. Cloudflare also announced general availability of AI Search on October 1, which is a different product. Web Search queries the open internet; AI Search is for your own indexed content. Verify the distinction in the docs for your use case.
A minimal search-grounded Worker
Here is a compact pattern that combines the search call with a model call. It keeps the search provider in config, caps results, and passes only short snippets to the model.
export default {
async fetch(request, env) {
const { question } = await request.json()
const search = await env.AI.websearch({
gatewayId: "default",
query: question,
provider: env.SEARCH_PROVIDER || "exa",
limit: 5,
})
const { results } = await search.json()
const context = results
.slice(0, 5)
.map((r, i) => `[${i + 1}] ${r.title} (${r.url})\n${r.description}`)
.join("\n\n")
const answer = await env.AI.run("@cf/meta/llama-3.1-8b-instruct", {
messages: [
{ role: "system", content: "Answer using only the numbered sources. Cite them as [n]. If the sources do not answer, say so." },
{ role: "user", content: `Sources:\n${context}\n\nQuestion: ${question}` },
],
})
return Response.json(answer)
},
}
Treat this as a sketch: the response field names and the model identifier should be checked against the current docs before you run it. The shape is what matters. Search first, trim, label sources, instruct the model to use only those sources, and let it say it does not know.
Why the sketch is shaped this way
Each choice in that Worker answers a failure you will meet in production. Reading the provider from an environment variable lets you switch without a deploy when one provider has an outage or a quality dip. Slicing to five results keeps the prompt small and the cost down. Numbering the sources lets the model cite them and lets you verify claims afterwards. And the instruction to answer only from the sources, and to say so when they do not answer, is the cheapest hallucination control you have. It does not eliminate errors, but it turns many confident guesses into honest refusals.
Controlling cost and quality
Search calls add up quickly when an agent loops. Four habits keep the bill predictable.
- Cap calls per task. Give each agent run a budget of searches, and stop when it is spent. An agent with no cap can issue dozens of near-identical queries.
- Cache by query. Identical or near-identical questions within a short window can reuse results. Even a five-minute cache removes a surprising amount of traffic.
- Keep
limitsmall. Five results is often enough. Fetching ten and discarding five pays for data you do not use. - Log provider per request. The gateway logs help you compare cost and success per provider over time, which is the data you need to change providers with confidence.
Quality has its own checks. Spot-check a sample of answers against their cited sources every week. If descriptions are thin, add a page-fetch step for the top result only. If answers cite stale pages, add a freshness filter on your side by preferring results with recent dates in the snippet.
Honest limitations
- I read Cloudflare's announcement and a launch tracker; I did not test the API.
- Beta status is unclear from the sources.
- No provider quality comparison is available yet.
- Pricing follows partner list prices, which can change.
Bottom line
Cloudflare's Web Search API gives agents live search behind AI Gateway with Ceramic, Exa or Linkup, billed at list prices and logged like model calls. It is a low-friction way to add grounding if you already use the gateway. Bake off the providers on your own queries, treat results as untrusted input, and watch for the native server tools.
Related on explainx.ai
- Cloudflare Sandbox SDK 1.0
- Cloudflare computer agent runtime
- Cloudflare Wallets for AI agent payments
- Cloudflare Artifacts: Git for agents
- Vercel AI Gateway and open models
- Indirect prompt injection in AI agents
- What is MCP?
- What is an agent harness?
Sources: Cloudflare Blog — Introducing Web Search API via AI Gateway · Cloudflare AI Gateway web search docs
Details reflect Cloudflare's October 2, 2026 announcement and may change; confirm status and pricing in the docs.
