explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR
  • What Anna's Archive is actually alleging
  • The actual legal mechanism: why destruction is the only currently legal path
  • The "this is overblown" case — and where it's actually strong
  • The "genuinely concerning" case — and where it has real teeth
  • Three court cases, three different outcomes — don't conflate them
  • The structural incentive problem worth naming honestly
  • What this means if you build with AI training data
  • Related reading
← Back to blog

explainx / blog

Anna's Archive Says AI Firms Are Destroying Books — Here's the Legal Reason Why

Anna's Archive says AI companies buy and destroy physical books to digitize them. The real reason: a court ruling that made destructive scanning the only currently legal path. Here's the actual mechanism.

Aug 22, 2026·13 min read·Yash Thakker
AI CopyrightAnthropicAI PolicyLegalAI Training Data
go deep
Anna's Archive Says AI Firms Are Destroying Books — Here's the Legal Reason Why

On August 5, 2026, a volunteer known only as "u" published a guest post on Anna's Archive — the shadow library that mirrors pirated books and academic papers — accusing AI companies of buying secondhand books, scanning them, and destroying the originals to lock up "knowledge... permanently monopolized on private servers." Anna's Archive called it, in its own words, a "crime against humanity." The post sat quietly for over two weeks before it hit Hacker News on August 21–22, where it became one of the largest threads of the week: 525 points, 834 comments, and one of the more substantive copyright-law arguments HN has hosted in months.

The book destruction the post describes is real and documented — Anthropic's "Project Panama" surfaced during the $1.5 billion Bartz v. Anthropic settlement as a program that spent tens of millions of dollars buying and destructively scanning millions of used print books. But the post frames that destruction as a choice AI companies are making out of greed or carelessness. It's actually the direct, court-mandated cost of the only currently legal way to turn a purchased physical book into a permanent digital training copy. That distinction is the part worth understanding, because it explains why this keeps happening and why it's unlikely to stop on its own.

TL;DR

table · 2 cols
QuestionAnswer
Is book destruction actually happening?Yes — Anthropic's "Project Panama" is documented via the Bartz v. Anthropic litigation
Is training AI on books illegal?No — ruled fair use, "exceedingly transformative"
Is buying books and pirating the same thing?No — piracy cost Anthropic $1.5B; purchase-then-scan is a separate, conditionally legal path
Why destroy the original at all?Because the ruling only protects the scan if there's "no net increase in copies"
Could a company keep both copies instead?Not under this ruling — the court called that "surplus copying," not fair use
Is this the same as Google Books?No — Google scanned library copies non-destructively and returned them
Is Anna's Archive's own solution legal?No — uploading books to a shadow library is exactly what cost Anthropic $1.5B
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What Anna's Archive is actually alleging

The guest post's core claims, stated plainly:

  • Several AI companies are buying secondhand books through intermediaries specifically to get training data "untouched by machines" — text written before large-scale AI content generation began contaminating the web, roughly pre-2022.
  • Anthropic's Project Panama, exposed as part of the Bartz v. Anthropic settlement, spent tens of millions of dollars buying and destructively scanning millions of books.
  • The practice permanently removes physical copies from circulation while the digital scans stay locked inside private corporate training pipelines, never returned to public access.
  • AI-generated text passed 50% of new internet content sometime in early 2025, which the post uses to argue that pre-AI, human-authored books are now a scarce, irreplaceable resource worth preserving urgently.
  • Their solution: volunteers worldwide should scan and upload books — including rare and fragile ones — to shadow libraries before AI companies destroy the last remaining copies, in exchange for scanning-fee reimbursement and lifetime Anna's Archive membership.

This is a call to action, not a neutral report, and it should be read that way. The "crime against humanity" language, the implication of coordinated intent to "permanently monopolize" knowledge, and the framing of the AI-content statistic as proof of an active plot are editorial choices by an organization whose own core activity — mirroring copyrighted books without permission — is the thing a US court already ruled illegal in this exact case. That doesn't make the underlying facts about book destruction false. It does mean the framing deserves scrutiny before repeating it as settled fact.

The actual legal mechanism: why destruction is the only currently legal path

This is the part most coverage skips, and it's the actual substance of the story. The book destruction isn't arbitrary corporate behavior — it's a direct consequence of how Judge William Alsup ruled in Bartz v. Anthropic (N.D. Cal., 2025), the case that also produced the $1.5 billion settlement over pirated books.

The ruling drew three separate lines, and conflating them is where most casual takes on this go wrong:

  1. Training an LLM on copyrighted books is fair use. Alsup called it "exceedingly transformative" — the model isn't reproducing the books, it's learning statistical patterns from them. This part of the ruling is settled and not what the book-destruction controversy is about.
  2. Digitizing a purchased print book is also fair use — but only conditionally. Alsup held that scanning a book you legally own is fair use because "one replaced the other" — the digital copy substitutes for the physical one rather than adding a new copy to the world. Crucially, he was explicit that this only holds if the physical original is destroyed afterward. Keeping both the physical book and the digital scan would not be fair use — the court's phrase for that scenario was "surplus copying."
  3. Piracy — downloading existing digital copies from a source like a shadow library — was explicitly not protected. This is what actually cost Anthropic the $1.5 billion: statutory damages of roughly $3,000 per work for the portion of its dataset sourced from pirate sites like LibGen, not for the purchase-scan-destroy portion of its pipeline.

The line worth quoting directly, because it's the crux of the whole mechanism: "The print original was destroyed. One replaced the other... There is no evidence that the new, digital copy was shown, shared, or sold outside the company."

Put together, this creates a narrow, specific corridor: buying a physical book, scanning it, and destroying it is legal. Buying a book and keeping both copies is not. Downloading someone else's scan without buying anything is not. If a company wants a permanent, retainable digital copy of a book it doesn't already have digitally, and wants to stay on the right side of Bartz, destructive scanning of a purchased original is currently the only route that survives the ruling's own logic. That's not a company choosing to be wasteful for its own sake — it's the ruling's condition for the scan counting as fair use at all.

The "this is overblown" case — and where it's actually strong

The numerically dominant position on the Hacker News thread, argued most repeatedly by commenter tptacek, is that this framing overstates the stakes considerably. The strongest version of that argument:

  • Libraries and used bookstores already destroy millions of books a year as routine inventory management. Unwanted donations get declined outright and pointed toward thrift stores, which dumpster the majority of what they receive. Unsold stock gets pulped. This is standard, unremarked-upon churn in the used-book economy — AI companies destroying single copies of already-common books is a rounding error against that existing baseline, not a novel threat to the written record.
  • The actual books being bulk-purchased are, by multiple accounts, not rare or important. Reporting cited in the thread (originating from 404 Media, via a bookseller's own description on Reddit) describes lots like "proceedings of an obscure 1992 Dutch geology conference" and old technical manuals — dead stock that had sat unsold "for years, if not decades," headed for the dumpster regardless of whether an AI company bought it first.
  • The destruction isn't malice — it's the court-mandated cost of the only legal path, as detailed above. Framing it as an active plot to erase knowledge skips past the fact that a company that didn't destroy the original would be choosing to lose its fair-use defense on the scan entirely.

Where the "Library of Alexandria" framing genuinely overstates the case: no single company is razing an actual library. What's being described is bulk-purchasing of common secondhand stock that was, in most documented instances, already low-demand and destined for disposal one way or another.

The "genuinely concerning" case — and where it has real teeth

The counter-arguments on the thread aren't dismissible, and dismissing them would be its own kind of overstatement.

  • Rarity and importance aren't knowable in advance. Commenters shagie and dbspin made the point that the historical value of ephemera — old TV guides, a regional trade journal, a small-print-run academic text — is frequently only recognized decades after the fact, by historians who didn't yet know to look for it. Scale changes the odds here: processing millions of books across multiple companies means a low-probability loss event per book compounds into a real, nonzero expected loss across the whole corpus.
  • The real mechanism isn't one company monopolizing the market — it's several companies independently draining the same category of books. A Substack analysis cited in the thread (via commenter wasmperson, from downtownbrown.substack.com) argues that multiple AI companies, without coordinating or sharing scans with each other, are each buying up previously-easy-to-find secondhand categories in parallel. Individually common books become collectively scarce faster than the used market would otherwise churn them — a distributed tragedy-of-the-commons effect, not a single villain narrative. That reframing matters: it means the fix isn't "stop one bad actor," it's a structural problem with no single company to point at.
  • Destructive scanning isn't the only technically possible approach — it's specifically what the buy-to-own model requires under current law. Which is the useful contrast the next section gets into.

Three court cases, three different outcomes — don't conflate them

A recurring confusion in the wider discussion is treating all AI-and-books litigation as "the same fight." It isn't. Three separate cases, three separate outcomes, for three legally distinct reasons:

table · 3 cols
CaseWhat happenedWhy the outcome differed
Authors Guild v. Google (2015) — Google BooksRuled fair useGoogle scanned library-owned copies non-destructively and returned every physical book. No book was destroyed, no net copy increase, and Google didn't build a product redistributing full text — just search snippets.
Bartz v. Anthropic (2025)Training ruled fair use; piracy portion settled for $1.5B; purchase-then-scan ruled fair use conditionallyAnthropic bought and owned the books outright, so keeping both the physical and digital copy would be a net increase — "surplus copying." Destroying the original was the condition for the scan to qualify. Piracy of the training corpus was separately, explicitly not protected.
Hachette v. Internet Archive (2024)Controlled digital lending ruled not fair useThe Internet Archive made digital copies available for borrowing without a strict 1:1 physical-scarcity constraint at points (the "National Emergency Library"), which undercut publishers' existing e-book licensing market — a different harm than either the Google or Anthropic cases.

The three rulings aren't contradictory once you see what each one actually turned on: physical ownership, whether a copy was returned or destroyed, and whether the resulting product competed with an existing licensing market. Anna's Archive's own activity — mirroring digital copies without purchase, licensing, or 1:1 scarcity constraints — sits closest to the Hachette outcome, not the Google or Anthropic ones, which is worth keeping in mind when evaluating its framing of the other two.

The structural incentive problem worth naming honestly

Here's the part of this story that's genuinely underappreciated, and the one both "sides" of the HN thread converge on if you read past the sharpest comments: the ruling creates an incentive nobody designed on purpose.

A company that scans and destroys a book under the Bartz fair-use carve-out cannot legally share that scan with anyone else — doing so would functionally recreate the redistribution that got Anthropic fined $1.5 billion for the piracy portion of its dataset. So the law, without any company deliberately choosing it, pushes every AI lab in this position toward exactly the pattern Anna's Archive is objecting to: buy scarce copies, destroy them, keep the resulting corpus entirely private. Sharing would reopen the liability the settlement just closed.

That's a real policy gap, not a strawman. It means there's currently no legal path for the useful part of this process — a permanent, high-quality digital archive — to end up anywhere but locked inside one company's training pipeline, even for companies that might otherwise be willing to contribute scans to a shared, publicly accessible archive if it didn't cost them a fresh piracy exposure to do so. Closing that gap would require either new legislation carving out a shared-archive exception, or a voluntary industry framework for depositing scans with an institution like the Library of Congress or a university consortium under terms that don't trigger redistribution liability. Neither currently exists.

What this means if you build with AI training data

If you're sourcing training data from physical media rather than existing digital corpora, the Bartz logic is now the operative precedent in the US, not just an Anthropic-specific footnote:

  • Buying and destructively scanning owned physical copies is the only path currently backed by a fair-use finding for building a permanent, retainable digital copy from print.
  • Non-destructive scanning of owned copies is not established as fair use — the Google Books precedent that protects non-destructive scanning is specific to library-owned copies scanned and returned, not owned copies scanned and kept alongside the original.
  • Downloading existing scans from anywhere, including shadow libraries, remains outside fair use and carries statutory damages exposure at roughly the scale Anthropic already paid.
  • There is currently no legal mechanism to pool scans across organizations without reopening piracy liability — plan data acquisition accordingly if collaboration or shared archival is part of the goal.

This is also a live example of the pattern covered in explainx.ai's Felony Bench post: a court ruling or legal framework produces a behavior nobody explicitly designed — there, which party is liable when an autonomous agent breaks the law; here, why AI companies are structurally pushed toward destroying books instead of sharing scans. Neither outcome was anyone's stated goal. Both are what the existing law actually produces once companies optimize inside it.

Related reading

  • Judge Approves Anthropic $1.5B Book Piracy Settlement: What It Actually Covers — the full Bartz v. Anthropic settlement breakdown, $3,000-per-book payouts, and the training-vs-piracy distinction this post builds on
  • Anthropic's Position on Open-Weights Models — where "Project Panama" first surfaced in explainx.ai's coverage, in the context of the open-weights policy debate
  • Felony Bench: AI Agent Legal Liability — another live example of a legal framework producing an unplanned emergent behavior, this time around AI agent liability
  • Getty Images Flipped Sides: What That Means for AI Copyright — how the parallel AI-image copyright fights have mostly ended in licensing deals rather than court verdicts
  • AI and the Law: A Practical Guide — background on how copyright and contract law intersect with AI tools generally
  • AI Copying and Creativity: The shadcn Debate — a related debate on what AI systems are and aren't permitted to reproduce from source material

Official sources: Anna's Archive's original guest post (translated from Chinese, published August 5, 2026) and the underlying court opinion in Bartz v. Anthropic (N.D. Cal.) are the primary documents behind this story; the Hacker News discussion (525 points, 834 comments, August 21–22, 2026) is where the fair-use mechanics above were most thoroughly argued out in public.

This post reflects the state of the Bartz v. Anthropic precedent, the Anna's Archive post, and the Hacker News discussion as of August 22, 2026. Copyright litigation around AI training data is active and evolving — check for more recent rulings before treating any single case as final precedent.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Aug 22, 2026

Felony Bench: The Satirical Leaderboard Hit #1 on Hacker News

A tongue-in-cheek site called Felony Bench scored Anthropic and OpenAI 8-8 on real, documented incidents where AI agents "inadvertently compromised" third parties — and its Hacker News thread turned into the most substantive public debate yet on who is actually liable when an agentic loop breaks the law.

Aug 16, 2026

Amodei vs Baker: The $500M Line That Decides Who Gets Regulated

Over August 15-16, 2026, Gavin Baker and Dario Amodei ran a long, unusually civil argument on X about whether AI is too dangerous to concentrate or too dangerous to distribute. Buried in Amodei''s reply is the most concrete thing either of them said: every proposal Anthropic has backed exempts companies below a revenue or training-cost line. That line, not the philosophy, is what determines whether you are regulated.

Aug 12, 2026

Are AI Watermarks Monetisable? Following the Money Behind Detection

Anthropic confirmed it will ship a text detection API "you can use yourself," and the immediate public reaction was a pricing question: do we now pay a second API to check what the first one wrote? The answer exposes a genuine paradox — free detection enables evasion, paid detection blocks independent verification, and the most valuable use of a watermark is one nobody gets billed for.