explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

[email protected]

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR — what's verifiable vs what's rhetoric
  • The concrete failures
  • The Internet Archive problem is more complicated than the essay says
  • The scraping arms race is degrading browsing for everyone
  • Wikipedia funds its own replacement
  • The strongest objection to the whole thesis
  • The publishing incentive question
  • What people are asking
  • The sovereignty argument
  • What to actually do
  • Related on explainx.ai
← Back to blog

explainx / blog

As AI Eats the Web, the Internet Is Losing Its Memory

Google AI summaries invent facts, FiveThirtyEight's archive was deleted, and Wikipedia is losing the traffic that funds it. The web's memory is failing — here's what publishers should do.

Aug 11, 2026·10 min read·Yash Thakker
AI SearchGEOInternet ArchivePublishingPolicy
go deep
As AI Eats the Web, the Internet Is Losing Its Memory

A Walrus essay by Vass Bednar published August 10, 2026 argued that the web's archival function is breaking down — and hit 287 points and 284 comments on Hacker News within hours. The thesis is worth taking seriously even where the headline overstates it: the problem isn't that search got worse. It's that the layer underneath search is getting thinner.

For anyone who publishes — which includes us — this is not a philosophical debate. It's a question about whether the thing you wrote in 2023 will be findable in 2029, and whether anyone will know you wrote it.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.


TL;DR — what's verifiable vs what's rhetoric

ClaimStatus
"Google Search is dying"Overstated. Search revenue is growing at an accelerating rate with expanding margins
AI Overviews produce confidently wrong answersTrue and now legally tested — a German court held Google liable
FiveThirtyEight's archive was deletedTrue. Disney removed nearly the entire archive after winding the site down
The Internet Archive is under real pressureTrue — lost the CDL suit, plus crawler blocking by news orgs
Wikipedia is losing the traffic that funds itDirectionally true; AI answers remove the click-through
Governments are moving off US searchTrue — France, European Parliament, Danish procurement
"The internet was ever a reliable cultural record"Contested — and the strongest objection to the whole framing

The concrete failures

The essay opens with a small, almost comic example: people missing sunsets because Google's AI summaries invented the times. One user in Colorado Springs had a projector set up outdoors and was told the sunset had already happened.

Trivial on its own — but it's the same failure mode that produced a real legal consequence. A German court recently held Google liable for false statements generated by its AI Overview feature, after the system wrongly linked two publishing companies to scammy business practices. The court's reasoning is the part that matters: because the search engine extracts and rewrites information in its own words, it isn't impartially pointing at the public record — it's authoring a new layer of content, and that carries editorial responsibility. Google is appealing.

That reasoning, if it holds, is the single most consequential development in this story. It converts AI summaries from "a feature" into "publishing," with everything that follows. It also isn't isolated — German regulators were already moving this direction, as we covered when ZAK examined AI Overviews and Perplexity under media law. Two independent German bodies have now landed on the same conclusion: answer engines are publishers, not conduits.

Meanwhile the record itself is thinning:

What was lostHow
FiveThirtyEight archiveDisney deleted nearly all of it once the site stopped being an active asset
Parts of the US ConstitutionBriefly vanished from the Library of Congress site via a coding error
Ordinary pagesLink rot, continuously
Ephemeral formatsInstagram Stories, WhatsApp statuses — never preserved in the first place

The FiveThirtyEight case is the clearest. A decade-plus of polling and elections analysis, removed by business decision. No villain, no controversy, no announcement worth noticing. Just an asset that stopped earning.

The Internet Archive problem is more complicated than the essay says

The Walrus frames publishers as having "successfully sued the Internet Archive over its digital lending program, calling it unauthorized copying." The top-voted correction on Hacker News is right to push back: the court determined the copying was unauthorized. That's a finding, not an allegation.

It's also not a case of publishers versus preservation in any simple sense. The Authors Guild, the National Writers Union, the European Writers Council, and the UK's Society of Authors all supported the suit. The NWU reports trying to open a dialogue as early as 2010 and being rebuffed.

Why this matters for the essay's own argument: the Wayback Machine — genuinely irreplaceable infrastructure — is now weakened partly by a strategic choice to keep litigating Controlled Digital Lending. Separately, news organizations began blocking Wayback crawlers out of fear that archived pages become an indirect pipeline to AI training data. Every AI-driven restriction on archiving compounds the exact memory loss the essay is warning about.

That's the real mechanism, and it's more uncomfortable than "AI is eating the web": the defensive reactions to AI are doing much of the damage.

The scraping arms race is degrading browsing for everyone

The practical version of this is already visible to anyone using the web. Sites now sit behind Cloudflare challenges, Anubis proof-of-work gates, and captchas, because scraper traffic is knocking them over.

One operator's summary from the thread: even behind Cloudflare, 90% of traffic is still scrapers. And the blocking mostly fails where it matters — well-behaved bots with honest user agents get blocked, while the problematic ones disguise themselves as Chrome and rotate through millions of residential proxy IPs. robots.txt is a request to the crawlers who were never the problem.

The choices site owners actually face: stay knocked offline, add user-hostile defenses, or unpublish. None of those improve the archive.

This connects to a shift we've tracked in AI traffic overtaking human traffic and the emerging economics of agent payments — if agents are most of your readers, "who pays for the page" becomes a protocol question, not a policy one.

Wikipedia funds its own replacement

The Wikipedia dynamic is the cleanest illustration of the whole problem. Search engines sent billions of visits; those visits produced both donations and new volunteers. AI systems ingest the content directly and answer without the click.

The encyclopedia becomes infrastructure for systems that don't sustain it. Nothing here requires malice — it's just a funding loop with the return path cut.

The strongest objection to the whole thesis

Worth stating fairly, because it's the best counterargument in the discussion: the premise that the web was ever a reliable cultural record is itself shaky.

Keeping anything online has always required somebody's continuous, unpaid effort — hosting renewed, domains maintained, enthusiasm sustained across decades. Things vanished constantly, long before AI. MUDs, IRC networks, forums, hobby sites, entire communities — most of that is gone, and blogs and social platforms killed more of it than LLMs have.

There's a second objection: not all of it is worth saving. The essay laments that Instagram Stories and WhatsApp statuses aren't preserved. For most of history, nobody tried to conserve every casual communication.

The counter to that — also good — is that historians have gotten enormous value from exactly this kind of debris. Graffiti in Pompeii tells us things Virgil doesn't. The difference now is that physical ephemera survived by default through neglect, while digital ephemera requires active effort to persist. We went from preserving a lot by accident to preserving almost nothing without intent.

The publishing incentive question

The sharpest line in the discussion: "AI will kill the internet because it is killing the incentive to make it."

Both sides are represented honestly:

PositionArgument
Stop publishingLLMs remix without attribution; no readers arrive; no collaboration requests; you train your replacement
Keep publishingIdeas are for sharing; frontier models demonstrably do retain authorship when asked; reaching people via an intermediary still reaches them

Several developers in the thread report making repos private or pulling old work offline. Whatever you think of that reasoning, it's a real behavioral change, and it removes exactly the independent technical writing that made the web useful.

Our position, stated plainly: explainx.ai publishes into this environment, so we have skin in it. We think the "stop publishing" reflex loses. What changes is where you publish and what you control — own the canonical URL, keep the archive, structure content so machines attribute it correctly, and stop treating a platform's servers as your memory. That's the practical core of generative engine optimization, and it's why AI slop and content quality is a durability issue rather than a taste issue.

For the marketing-side view, see GEO for marketers.

What people are asking

"Is Google actually worse, or is the web worse?" — Mostly the second, but not entirely. SEO spam and AI content farms genuinely degraded the corpus. But Google monetizes much of that spam through its own ad network, so the "we're just victims of SEO" framing doesn't fully hold. The "Crawled, currently not indexed" complaint from webmasters is real: a short, genuinely useful post with no backlinks and no domain authority now often fails to get indexed at all.

"Are the alternatives better?" — Partially. Kagi is the most cited, and its paid model removes the advertiser-as-customer inversion. It also can't fix the underlying corpus, and it depends on other search APIs. Brave and DuckDuckGo have their own tradeoffs. Nothing available replaces a healthy index.

"Is search revenue actually falling?" — No, and this is where "dying" fails. Search revenue is growing at an accelerating rate with expanding operating margins. Inferring business collapse from your own switch to an LLM is a common error. The degradation is real; the death is not.

"Is AI search monetization coming?" — Yes. OpenAI is already running ads in ChatGPT. There are no ad blockers for LLM answers. The current experience is the best it will be.

The sovereignty argument

The essay's most actionable section is about governments treating information retrieval as infrastructure:

  • France adopted Qwant across its national assembly and armed forces ministry starting in late 2018; the European Parliament — 720 members plus thousands of staff — followed.
  • France requires civil servants to use Tchap instead of WhatsApp or Signal.
  • European governments have replaced Microsoft products with open-source alternatives across hundreds of thousands of workstations.
  • Denmark's digitalization minister, Caroline Stage Olsen: "Far too much public digital infrastructure is today tied up with very few foreign suppliers. This makes us vulnerable."

The analogy the essay lands: roads aren't governed by ride-sharing companies, and payment apps don't set monetary policy. Search has been treated as a free consumer service rather than infrastructure, and volunteerism alone — Wikipedia, the Wayback Machine — cannot hold up the public record.

What to actually do

If you publish anything you want to exist in five years:

  1. Own the canonical URL. Platform-hosted content is borrowed shelf space; FiveThirtyEight had a corporate parent and still lost the archive.
  2. Keep your own copies. Local exports of anything you'd be upset to lose. This is now a normal practice, not paranoia.
  3. Structure for machine attribution. Schema, clear authorship, stable canonicals. If the AI layer is the reader, make it a reader that cites you.
  4. Assume the click is optional. Build value that survives being summarized — depth, data, and specificity that a three-sentence answer can't carry.
  5. Support the archives. The Internet Archive and Wikipedia are load-bearing and underfunded.

The uncomfortable truth in the essay's last line: even that article won't be online forever. Neither will this one, unless someone keeps deciding it should be.

Related on explainx.ai

  • What is GEO? Generative engine optimization — publishing for a web where AI is the reader
  • GEO for marketers — the practical playbook
  • AI slop and content quality — what's polluting the corpus
  • AI traffic overtakes human traffic — the crawler economics
  • OpenAI brings ads to ChatGPT — monetization arrives at the answer layer
  • Cloudflare Wallets and agent payments — paying for content when agents are the audience
  • Germany's ZAK on AI Overviews and Perplexity — answer engines treated as publishers under media law

Source: "Google Search Is Dying. What Comes Next Is Worse" by Vass Bednar, The Walrus, August 10, 2026


Accurate as of August 11, 2026. Court rulings described are subject to appeal — the German AI Overview decision is being appealed by Google. Search revenue figures reflect the most recent public reporting at time of writing.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Apr 10, 2026

The seo-geo agent skill: SEO plus GEO for Google, Bing, and AI answer engines

seo-geo is an agent-installable playbook for technical audits, keyword research, structured data, and GEO tactics tuned for ChatGPT-style citation surfaces — live on explainx.ai with sources on GitHub.

Aug 11, 2026

ChatGPT Books Your Table Now: Inside the OpenTable, Resy, and Yelp Integrations

On August 10, 2026, Yelp brought Reservations and Waitlist into ChatGPT for thousands of US and Canada restaurants, alongside Resy in the US and OpenTable globally. explainx.ai breaks down the mechanism, the recommend-vs-transact jump, the local GEO implications, and what surface you need to expose to be bookable by an assistant.

Aug 11, 2026

Zuckerberg's "Abundance of Jobs" Claim, Checked Against the Data

Zuckerberg's August 10 essay argues AI brings an abundance of jobs — world builders, personal biologists, one-person studios. Prediction markets broadly agree with the near-term calm. explainx.ai checks the claim against actual labor data and finds the argument sound on the endpoint and silent on the part that hurts.