Anthropic’s Book Scanning Strategy Could Set a Pattern
Anthropic is buying printed books, scanning them for training, and destroying the originals to argue the copies are legally defensible.

Anthropic is buying printed books, scanning them for training, and destroying the originals.
Anthropic is using a legal strategy that could shape how AI companies collect training data from books. The idea is simple: buy a physical copy, scan it into text, destroy the paper version, and argue that the digital file replaces the original without expanding the total number of copies in circulation.
What Anthropic is trying to defend
Get the latest AI news in your inbox
Weekly picks of model releases, tools, and deep dives — no spam, unsubscribe anytime.
No spam. Unsubscribe at any time.
The theory leans on two familiar parts of U.S. copyright law: the first sale doctrine and fair use. First sale limits what copyright holders can control after a lawful sale of a physical item. Fair use can protect certain transformative uses, especially when the new use does something different from the original work.

That matters because AI training often happens in a legal gray zone. A model does not store books the way a library does, but training does involve making copies during ingestion. Courts then have to decide whether that copying is the kind of technical, transformative use copyright law allows.
Anthropic’s argument is that a scanned book is a substitute for the destroyed paper copy, not an extra unauthorized copy added to the market. If a court accepts that framing, the company gets a cleaner defense than if it had simply scraped text from the web without buying the books first.
- Physical book purchased legally
- Book scanned into digital text
- Paper copy destroyed after digitization
- Digital file used for model training
Why this matters beyond one lawsuit
If this legal theory holds, other AI labs may copy the same workflow. Buying books in bulk is slower and more expensive than downloading text from the internet, but it creates a paper trail that looks far better in court. For companies already facing copyright claims, that paper trail can be worth the extra cost.
The bigger issue is what this says about training data norms. A company that can show it bought and destroyed physical books may argue it acted more carefully than a company that scraped millions of pages from online libraries, forums, and publisher sites. That does not settle the copyright question, but it changes the optics and the litigation strategy.
There is also a practical limit. Book buying and scanning only works for a slice of the data AI companies want. Large models need code, news, academic papers, manuals, images, and conversation data. A legal win for books would not automatically solve the broader problem of how AI firms obtain training material across formats.
“A machine learning model is trained on data that is not itself expressive in the same way as the source material.” — OpenAI, in its copyright and AI discussion
How this compares with other AI training fights
Anthropic is not the only company dealing with this issue. OpenAI has faced lawsuits over training data, and Meta has also been pulled into copyright disputes around model training. The common thread is that courts are being asked to decide whether mass copying for model training is a protected technical process or an infringement problem.

There are some useful contrasts here:
- Buying and scanning books creates a lawful purchase record.
- Web scraping can involve material with unclear permissions.
- Licensed datasets reduce legal risk but raise cost and access limits.
- Public-domain books avoid many copyright issues but cover only older texts.
That comparison explains why Anthropic’s approach is getting attention. It is a middle path between open web scraping and fully licensed data deals. It does not remove legal exposure, but it may reduce it enough to matter in court.
It also shows how AI firms are adapting to litigation pressure. Instead of defending every data source as fair use, some companies may start building training pipelines that look more like archival digitization. That changes procurement, compliance, and storage policies all at once.
The likely effect on AI training policy
If courts accept the logic, AI companies may start buying more physical books, scanning them internally, and destroying the originals as standard practice. Publishers would probably push back by arguing that the digital copy still competes with licensed e-books and other editions, even if the paper original is gone.
That fight could spill into policy and contracts. Publishers may demand clearer licensing terms, stricter audit rules, or explicit limits on digitization for model training. AI companies, meanwhile, may treat physical acquisition as one more compliance step rather than a full solution.
The real question is whether courts treat this as a narrow workaround or a durable rule. If the answer is yes, expect more AI firms to build training pipelines around purchased books and documented destruction. If the answer is no, the industry will need licenses, settlements, or new data standards that can survive in court.
For now, Anthropic’s move is less a final answer than a test case. It asks a simple question with expensive consequences: if a company buys a book, scans it, and destroys it, has it copied the work in a way copyright law forbids, or has it simply changed the format of a lawful purchase?
// Related Articles
- [IND]
Millions Raised for Zhipu-style Social World Model
- [IND]
Huang’s open-letter playbook for open-weight AI
- [IND]
32 firms back open-weight AI in DC letter
- [IND]
Huang usa il suo primo post su X per difendere l’IA aperta
- [IND]
Black Duck’s Coverity gets better at AI-era triage
- [IND]
Anthropic’s Opus 5 makes the AI race cheaper