explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR — what people are asking
  • The measurement problem in one sentence
  • Why dashboards oversell — six mechanisms
  • What AI visibility can honestly tell you
  • Approaches that measure more honestly
  • Audit checklist before you buy or act
  • What to do Monday morning
  • Connection to Fable "inner voice" and model stacks
  • Define the observation before counting it
  • Keep a stable panel when measuring a change
  • Connect content experiments to a user outcome
  • Related Reading
← Back to blog

explainx / blog

Can You Trust AI Visibility Scores? Why AEO Dashboards Oversell Precision

AEO, GEO, AI visibility, SEO, Measurement, Marketing

AI visibility tools promise rank #4 and 17% share of voice — but ChatGPT and Claude answers are noisy, geographic, and nondeterministic. What AEO metrics can honestly tell you, Canonry's critique, and better measurement approaches.

Jul 3, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
Can You Trust AI Visibility Scores? Why AEO Dashboards Oversell Precision

Can you trust your AI visibility score? A July 2026 Hacker News thread on Canonry's essay — "Every AI Visibility Tool Is Lying to You" — hit a nerve. Not because brands should ignore ChatGPT, Claude, Gemini, and Perplexity citations — but because dashboards turn messy, personalized, nondeterministic answers into fake precision: you are rank #4, 17% share of voice, +2 positions this week.

One HN comment captured the implementation dread: "I'm implementing changes suggested by a stochastic model based on a limited set of searches on other stochastic models."

This post answers what AEO / GEO measurement can honestly tell you — and lists practical approaches (Canonry's distribution model among them, not the only one).


Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR — what people are asking

table · 2 cols
QuestionHonest answer
Are visibility tools useless?No — good for directional gaps (invisible on category prompts, missing in NYC)
Is "rank #4 in ChatGPT" real?Over-precise — one sample from a distribution unless proven with variance
Scrape vs API?Different instruments — neither equals "what every customer sees" without disclosure
Why three tools disagree?Prompt sets + scoring formulas manufacture different headline numbers
Local businesses?Global rank is meaningless — geography must be explicit
What works better?Repeated runs + raw evidence + GEO fundamentals

AI visibility and AEO measurement tools explained — one sample dot pulled from a wide scatter distribution into a falsely precise dashboard gauge

The measurement problem in one sentence

The same words often produce different brand lists on the next run — SparkToro/Gumshoe volunteer studies and production temperature-zero instability (Thinking Machines Lab on batching variance) both show this. A point estimate without a distribution is decoration.


Why dashboards oversell — six mechanisms

1. Frontend scrape = one synthetic user

Scraping ChatGPT or Claude sounds like the real product. It is one account, geography, memory state, subscription tier, and IP story. A buyer in Brooklyn and a datacenter browser asking "best CRM for seed-stage startup" are different experiments.

2. API ≠ consumer app

API calls are cheaper, repeatable, auditable — but may lack consumer memory, routing, shopping modules, and UI-specific retrieval. OpenAI requires explicit web search tools; Gemini has its own grounding config. API measurement is valid when labeled as API, not "what the app showed your buyer."

3. Prompt sets manufacture the score

Vendors track 100–1,000 prompts (Profound's own design guidance). Change the list — "best AEO agency NYC" vs "SEO agency" — and "visibility" changes. Same evidence, three scoring formulas (mention SOV, position-weighted, citation-based) can yield 20% vs 16.8% vs 31.4% on Digital Applied's framework example.

4. Geography breaks global leaderboards

"Best roofing company near me" is local. A single global number without city, proxy, or explicit geo context is marketing math.

5. Model drift moves the goalposts

Chen, Zaharia, and Zou documented GPT-4 behavior shifts under the same public name (e.g. prime accuracy 84% → 51% March–June 2023). OpenAI rolled back a GPT-4o update for being too agreeable (April 2025). Your "+2 this week" may be the model, not your blog post.

6. Recommendations stacked on recommendations

Tool runs prompts → another model summarizes into prose → your team ships portal copy. Probability layered on probability — the HN engineer's complaint is structurally fair.


What AI visibility can honestly tell you

Directional, probabilistic findings are useful:

  • Invisible on commercial prompts buyers ask
  • Strong on branded prompts, weak on category prompts
  • Competitor cited more often with source links
  • Visible in New York, blank in Los Angeles
  • Schema/citation change correlates with more mentions over repeated runs

Fake precision:

  • "You are rank #4"
  • "17% AI share of voice" (single number, no interval)
  • "This week's lift was caused by last week's post"
  • "This screenshot is what customers see"

Approaches that measure more honestly

Not one vendor — a stack:

table · 4 cols
ApproachWhat it measuresStrengthWeakness
Vendor dashboard (Profound, etc.)Sampled prompt panelFast baselineOften hides methodology
CanonryRepeated API observations + evidence; local runs for geoDistribution mindset; auditable runsCosts more; still not every user session
DIY prompt panelYour prompts × providers × cities × N runsFull controlLabor; no pretty UI
GEO content workCitations, schema, entity consistencyCompounds over monthsNot a weekly rank chart
seo-geo agent checksPrinceton GEO methods in contentImproves cite-worthinessNot competitive monitoring
Classic analytics + branded searchTraffic, conversions, Search ConsoleGrounded in outcomesMisses dark-chatGPT referrals

Canonry's specific pitch: treat visibility as a distribution — multiple runs, multiple providers, explicit geolocation on APIs, stored evidence. Local-first execution for agencies serving Chicago HVAC or Brooklyn hospitality — run probes from machines in-market instead of only vendor cloud regions. That addresses one scrape bias; it does not eliminate model nondeterminism.


Audit checklist before you buy or act

Ask any tool (or your internal script):

  1. Frontend scrape, API, or both?
  2. Whose account, tier, memory, location?
  3. How many runs per prompt → one number?
  4. Variance or confidence intervals reported?
  5. Full prompt list and weights?
  6. Scoring formula (mention vs citation vs position)?
  7. Raw answers and cited domains retained?
  8. Model version logged to separate drift from your changes?

Without answers, do not rewrite production docs on a single score.


What to do Monday morning

  1. Pick 20–50 real buyer prompts — not only vanity category terms
  2. Run each 5–10 times across ChatGPT, Claude, Gemini, Perplexity
  3. Store JSON — brands mentioned, citations, timestamp, city
  4. Plot presence rate, not rank
  5. Fix GEO basics — FAQ schema, stats, linked sources (citation is the new ranking)
  6. Review recommendations as experiments — A/B docs, measure signup and support tickets

Connection to Fable "inner voice" and model stacks

Separate July 2026 thread — Fable leaking reasoning traces — reinforces the same theme: you are watching one layer of a stochastic stack. Visibility scores read model outputs; implementation guides read another model's summary of those outputs. Show your work or stay skeptical.


Define the observation before counting it

A brand mention, a cited domain, and a recommendation are different events. An answer can mention your company while advising against it. It can cite a page without naming the brand. Decide which event matters to the question you are investigating and retain enough raw output to classify it correctly.

For an illustrative panel, suppose a brand appears in four of ten repeated answers to one prompt. The observed mention rate for that prompt is forty percent. That does not mean forty percent of all potential buyers will see the brand. The sample says something about the tested prompt and conditions; its broader meaning depends on how representative those conditions are.

Keep negative mentions and unsupported recommendations visible. A dashboard that collapses them into “visibility” can make harmful exposure look like progress. Where automated extraction classifies answers, inspect a sample by hand and record disagreements instead of treating the extraction model as ground truth.

Keep a stable panel when measuring a change

Save the original prompt list before editing your pages. Repeat that same panel after the change under comparable conditions, and record provider or model changes that occurred in between. Otherwise, a different prompt mix or changed search surface can masquerade as an effect of your content update.

Separate exploratory prompts from the tracked panel. Exploration helps discover a new buyer question. Adding that question to a future panel is reasonable, but its appearance should not silently change the historical denominator. Version the panel so a reader can explain why a total changed.

For a local business, specify the place in the prompt when location is part of the tested question. Record whether you tested an API request, a browser account, or another supported surface. Avoid describing that sample as the experience of everyone in the city unless your measurement actually supports that claim.

Connect content experiments to a user outcome

Before accepting a vendor suggestion, identify the underlying page problem. Is the business name inconsistent? Is a product capability described ambiguously? Is a useful source missing? Fixing those issues can help readers directly, regardless of whether a visibility metric moves next week.

Choose an outcome you can observe alongside answer samples: relevant referral traffic, qualified inquiries, fewer support misunderstandings, or an improved conversion path. Each has limits, and none captures every AI-assisted discovery. Looking at several signals is still more useful than attributing a dashboard fluctuation to a single copy edit.

Retain unsuccessful experiments. If a substantial rewrite did not change the tested answers, that result prevents the team from repeating the same intervention under a new name. Honest AEO measurement is a record of what was sampled, what changed, and what remained uncertain. It should sharpen a content decision rather than supply a decorative rank.

Related Reading

  • Armature study: what Claude Code, Codex, and Cursor actually pick — a vendor-run measurement study on agent tool choice, with the same "who's measuring, and why" caveat this post raises
  • Cloudflare Monetization Gateway — paywalling agent traffic with x402
  • What Is SEO & GEO?
  • seo-geo Agent Skill for AI Search
  • What Is AI Slop — Quality vs Volume
  • Fable Inner Voice Leak — Stochastic Stacks
  • Can LLMs Watch Video? — Distribution Sampling
  • Canonry — original essay

Analysis informed by Canonry's June 30, 2026 essay and HN discussion July 2026. Vendor features change — verify methodology on any tool before budget decisions.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Jul 9, 2026

GEO for Marketers: How to Get Cited by ChatGPT, Perplexity, and AI Overviews in 2026

Your customers are asking ChatGPT and Perplexity for recommendations before they ever see a search results page. Here's how marketing teams adapt content, structure, and workflow so AI answer engines cite them instead of a competitor.

May 26, 2026

What is SEO-GEO? Generative Engine Optimization explained (2026)

Generative Engine Optimization (GEO) is how you get cited—not ranked—in ChatGPT, Perplexity, Gemini, and AI Overviews. Here is what SEO-GEO means, why it matters now, and how to apply it without chasing hacks.

Apr 10, 2026

The seo-geo agent skill: SEO plus GEO for Google, Bing, and AI answer engines

seo-geo is an agent-installable playbook for technical audits, keyword research, structured data, and GEO tactics tuned for ChatGPT-style citation surfaces — live on explainx.ai with sources on GitHub.