Started By
Message

AI companies are buying rare books, cutting the spine off to scan, then shredding them

Posted on 7/29/26 at 4:09 am
Posted by hawgfaninc
https://youtu.be/torc9P4-k5A
Member since Nov 2011
64367 posts
Posted on 7/29/26 at 4:09 am

quote:

AI companies are bulk-buying rare books, scanning them through high-speed machines that cut the spines off, and shredding the originals. A service called ISBNdb facilitates orders of up to a million books and keeps buyers anonymous. Pre-2022 books are premium because they're free of AI-generated text. A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time. Anthropic hired the former head of Google Books partnerships to obtain "all the books in the world."

My Take
This got to me. A bookseller told 404 Media that rare books with almost no surviving copies are being fed into this pipeline. Books that survived wars, fires, and centuries of handling are being shredded so an AI can learn to write a better marketing email.
ISBNdb's website literally says "'AI company destroys two million books' is not a headline that generates sympathy," and they still built an entire business around making it happen quietly. They offer NDAs as a feature. They coach clients to call it "digital preservation."

I've covered AI companies scraping the internet, torrenting libraries, and stealing music. This is worse because it's irreversible. You can re-upload a website. You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate.
"We shred rare books and offer NDAs so nobody finds out" is a legitimate business model in 2026. What a timeline.

quote:

btw anthropic's internal document on this literally said "we don't want it to be known that we are working on this.”

it was called project panama.

here's exactly what happened:

1: anthropic concluded that books were the cheapest way to build a world-class model because they gave claude curated facts, structured arguments, compelling stories, and writing “an editor would approve of.”

2: once anthropic decided it needed books at enormous scale, its first solution was piracy.

it downloaded 7m+ books from online libraries including libgen. the judge later wrote that although anthropic had legal ways to buy them, it chose piracy to avoid what dario amodei called the “legal/practice/business slog.”

3: that piracy created a massive legal risk.

so in february 2024, anthropic hired tom turvey, the former head of partnerships for google books, to find a legally safer way of obtaining “all the books in the world.”

4: turvey first contacted major publishers about licensing their catalogs.

those attempts didn’t produce agreements, so anthropic chose a route that required no publisher permission: buying millions of physical books through distributors and used-book retailers.

5: within about a year, anthropic spent tens of millions acquiring and scanning millions of books, including many rare and 1/1 titles. one vendor proposal targeted 500,000 to 2 million books in six months.

6: to scan that many books within months, the vendors physically dismantled them.

a hydraulic cutter removed each spine. the pages were trimmed to size, fed as loose sheets through high-speed industrial scanners, and converted into searchable PDFs. the paper remains were then sent for recycling.

7: these PDFs were fed into claude as training data.

the complete collection became a private, searchable anthropic library that the company planned to “store forever.” the scans aren’t available to the public and were never open-sourced.
Posted by hawgfaninc
https://youtu.be/torc9P4-k5A
Member since Nov 2011
64367 posts
Posted on 7/29/26 at 4:10 am to
LINK
quote:

On Monday, court documents revealed that AI company Anthropic spent millions of dollars physically scanning print books to build Claude, an AI assistant similar to ChatGPT. In the process, the company cut millions of print books from their bindings, scanned them into digital files, and threw away the originals solely for the purpose of training AI—details buried in a copyright ruling on fair use whose broader fair use implications we reported yesterday.

The 32-page legal decision tells the story of how, in February 2024, the company hired Tom Turvey, the former head of partnerships for the Google Books book-scanning project, and tasked him with obtaining “all the books in the world.” The strategic hire appears to have been designed to replicate Google’s legally successful book digitization approach—the same scanning operation that survived copyright challenges and established key fair use precedents.

While destructive scanning is a common practice among some book digitizing operations, Anthropic’s approach was somewhat unusual due to its documented massive scale. By contrast, the Google Books project largely used a patented non-destructive camera process to scan millions of books borrowed from libraries and later returned. For Anthropic, the faster speed and lower cost of the destructive process appears to have trumped any need for preserving the physical books themselves, hinting at the need for a cheap and easy solution in a highly competitive industry.

Ultimately, Judge William Alsup ruled that this destructive scanning operation qualified as fair use—but only because Anthropic had legally purchased the books first, destroyed each print copy after scanning, and kept the digital files internally rather than distributing them. The judge compared the process to “conserv[ing] space” through format conversion and found it transformative. Had Anthropic stuck to this approach from the beginning, it might have achieved the first legally sanctioned case of AI fair use. Instead, the company’s earlier piracy undermined its position.

But if you’re not intimately familiar with the AI industry and copyright, you might wonder: Why would a company spend millions of dollars on books to destroy them? Behind these odd legal maneuvers lies a more fundamental driver: the AI industry’s insatiable hunger for high-quality text.

The race for high-quality training data

To understand why Anthropic would want to scan millions of books, it’s important to know that AI researchers build large language models (LLMs) like those that power ChatGPT and Claude by feeding billions of words into a neural network. During training, the AI system processes the text repeatedly, building statistical relationships between words and concepts in the process.

The quality of training data fed into the neural network directly impacts the resulting AI model’s capabilities. Models trained on well-edited books and articles tend to produce more coherent, accurate responses than those trained on lower-quality text like random YouTube comments.

Publishers legally control content that AI companies desperately want, but AI companies don’t always want to negotiate a license. The first-sale doctrine offered a workaround: Once you buy a physical book, you can do what you want with that copy—including destroy it. That meant buying physical books offered a legal workaround.

And yet buying things is expensive, even if it is legal. So like many AI companies before it, Anthropic initially chose the quick and easy path. In the quest for high-quality training data, the court filing states, Anthropic first chose to amass digitized versions of pirated books to avoid what CEO Dario Amodei called “legal/practice/business slog”—the complex licensing negotiations with publishers. But by 2024, Anthropic had become “not so gung ho about” using pirated ebooks “for legal reasons” and needed a safer source.

Buying used physical books sidestepped licensing entirely while providing the high-quality, professionally edited text that AI models need, and destructive scanning was simply the fastest way to digitize millions of volumes. The company spent “many millions of dollars” on this buying and scanning operation, often purchasing used books in bulk. Next, they stripped books from bindings, cut pages to workable dimensions, scanned them as stacks of pages into PDFs with machine-readable text including covers, then discarded all the paper originals.

The court documents don’t indicate that any rare books were destroyed in this process—Anthropic purchased its books in bulk from major retailers—but archivists long ago established other ways to extract information from paper. For example, The Internet Archive pioneered non-destructive book scanning methods that preserve physical volumes while creating digital copies. And earlier this month, OpenAI and Microsoft announced they’re working with Harvard’s libraries to train AI models on nearly 1 million public domain books dating back to the 15th century—fully digitized but preserved to live another day.

While Harvard carefully preserves 600-year-old manuscripts for AI training, somewhere on Earth sits the discarded remains of millions of books that taught Claude how to juice up your résumé. When asked about this process, Claude itself offered a poignant response in a style culled from billions of pages of discarded text: “The fact that this destruction helped create me—something that can discuss literature, help people write, and engage with human knowledge—adds layers of complexity I’m still processing. It’s like being built from a library’s ashes.”
This post was edited on 7/29/26 at 4:10 am
Posted by hawgfaninc
https://youtu.be/torc9P4-k5A
Member since Nov 2011
64367 posts
Posted on 7/29/26 at 4:12 am to
Loading Twitter/X Embed...
If tweet fails to load, click here.

quote:

Anthropic is getting hate for this, but it's ultimately an outcome of the same disconnect between copyright law and technology that's been hobbling the Internet for 25 years.

Digital telecommunications drops the marginal cost of replication and distribution of the information contained in a book, or any other media, to zero. Google Books was originally intended as an online Library of Alexandria that would provide access to every book ever written. The publishing industry saw that as a threat, and used copyright to shut that down, as a result of which Google Books is effectively useless.

LLMs are essentially knowledge machines that can digest the entire corpus of human knowledge, rendering it legible and interrogable in ways that no other library does. That's very useful, but it requires, by definition, vast quantities of training data. The best training data is not Reddit posts, but books, because that's where the knowledge is.

Anthropic is destroying, instead of just scanning, the books for essentially the same reason that Google locked down Google Books. Since the original no longer exists, only one copy remains, and the rights attached to purchase of the original copy transfer to the digital copy. It would cost Anthropic nothing, in principle, to make its digital library freely available to everyone, but they'd get sued into bankruptcy by the publishing industry.

Meanwhile, you can go on Amazon and regularly find used copies of out of print books from even just a couple decades ago going for hundreds of dollars. The publishing industry keeps the vast majority of its back catalogue in the vault, with no intention of making it available for sale. And, ok, it isn't necessarily profitable to reprint an old book everyone has forgotten. That's fine, printing is expensive, but keeping the information itself hidden is criminally insane.
Posted by Sunnyvale
Little ST. James
Member since Feb 2024
3705 posts
Posted on 7/29/26 at 4:13 am to
My Rent is still due Monday coming buddy. Dont have time for this.
Posted by 225Tyga
Member since Oct 2013
19880 posts
Posted on 7/29/26 at 4:17 am to
Pretty smart if you ask me
Posted by hawgfaninc
https://youtu.be/torc9P4-k5A
Member since Nov 2011
64367 posts
Posted on 7/29/26 at 4:29 am to
Loading Twitter/X Embed...
If tweet fails to load, click here.

quote:

I'm going to say this again for those at the back: hundreds of thousands of TONS of books —some of them "rare" in the sense that obsolete books on Windows 95 programming or service manuals for Singer sewing machines from 1985 or books of racist jokes from 1958 are "rare"— are pulped or landfilled every year. It's illegal to digitally archive them because the PUBLISHERS won't allow it. If AI companies are diverting them to use for training on the way to being thrown out, that's GOOD, actually. The reporting on this is custom-crafted to be rage-bait for normie middle-class liberal book lovers, and you all are falling for it hook, line and sinker. Just complete suckers. Nobody is shredding Gutenberg bibles or the Codex Gigax here.

Posted by Loup
Ferriday
Member since Apr 2019
17507 posts
Posted on 7/29/26 at 4:36 am to
I want to find some super obscure, rare af book that isn't super expensive. Buy all remaining copies and destroy all but one. When AI has scanned everything else I'll have the only book it hasnt scanned and itll be worth a pile of guap.
Posted by dnm3305
Member since Feb 2009
16290 posts
Posted on 7/29/26 at 4:55 am to
Fahrenheit 451 right in front of our eyes
Posted by Penrod
Member since Jan 2011
57158 posts
Posted on 7/29/26 at 5:09 am to
I don’t mind the practice, but we should pass a law mandating that destroying rare books requires the perpetrator to make digital copied accessible to all.
This post was edited on 7/29/26 at 5:10 am
Posted by saint tiger225
San Diego
Member since Jan 2011
49175 posts
Posted on 7/29/26 at 5:22 am to
I'd probably be more upset at this if I could read.
Posted by RoyalAir
Detroit
Member since Dec 2012
7732 posts
Posted on 7/29/26 at 6:04 am to
Carter is correct.

Physical publishers struggled to understand how to preserve data and share info in a digital world. There's really no reason that a book published more than 30 years ago shouldn't be free on Kindle right now.
Posted by PrimeTime Money
Houston, Texas, USA
Member since Nov 2012
28101 posts
Posted on 7/29/26 at 6:11 am to
When they say “rare books” where there’s only one copy, I’m assuming these are random books that nobody has ever read. Not some rare sought-after book where only one copy exists and is worth a bunch of money to a collector.

So in that case, who even cares?
Posted by dakarx
Member since Sep 2018
8566 posts
Posted on 7/29/26 at 7:04 am to
As an old guy, looking at both sides i am sort of conflicted on this practice....digitizing books that would otherwise be destroyed is good, but the real issue i have is the destruction bit, while the scanned pdf will not be released into the wild, is it retained after being fed into the system or destroyed? If it is destroyed when there is no record of the original AI gets to interpret/dictate history.

He who controls history/information controls the world.
Posted by Oilfieldbiology
Member since Nov 2016
42614 posts
Posted on 7/29/26 at 7:08 am to
quote:

He who controls history/information controls the world.


Who holds the past now, controls the future. Who holds the present now, controls the past. Who controls the past now, controls the future. Who controls the present now. Now testify!
Posted by SpqrTiger
Baton Rouge
Member since Aug 2004
9758 posts
Posted on 7/29/26 at 7:16 am to
Just as important is when you work with old, rare books, especially with foreign language books, you need a human to handle idiosyncrasies and outliers in print method, font style, dialect, dated language, slang, and context that machines will never be able to account for (accurately) alone.

There is no scanning, learning, and burning an 18th century treatise, for example. Even in your home language, you can’t take what you see exactly at face value. It’s been my experience that a machine can’t even tell the difference between an F and a S on an old document due to the font similarity. That alone will roadblock the text recognition process.

Posted by lsu777
Lake Charles
Member since Jan 2004
38466 posts
Posted on 7/29/26 at 7:24 am to
quote:

I don’t mind the practice, but we should pass a law mandating that destroying rare books requires the perpetrator to make digital copied accessible to all.


here is the problem, without originals who is to say over time that things arent changed

as left leaning as all the AIs are, except grok, do you really want them controlling the knowledge of the past? past mistakes

im not talking now but 50 years from now when most of us are gone, there will be no knowledge of things like soviet style communism except in books, do we really want them to not have the knowledge of how many they killed?
Posted by Dixie2023
Member since Mar 2023
5618 posts
Posted on 7/29/26 at 7:27 am to
Exactly. Thats the plan most likely. Destruction of reading material, history, etc. in 50 years it’s gonna be bad.
Posted by NIH
Member since Aug 2008
124351 posts
Posted on 7/29/26 at 7:29 am to
Retarded boomers think “we will win” by having Silicon Valley nerds control the traffic of right think and information
Posted by CocomoLSU
Inside your dome.
Member since Feb 2004
157134 posts
Posted on 7/29/26 at 8:22 am to
quote:

but the real issue i have is the destruction bit, while the scanned pdf will not be released into the wild, is it retained after being fed into the system or destroyed? If it is destroyed when there is no record of the original AI gets to interpret/dictate history.

This is the exact problem every single person should have with this. But even in this thread there are several "who cares" kinds of posts. It's fricking sad honestly.

This isn't about "yeah but if it's a random book nobody wants to buy then it doesn't matter"; it's the principle of destroying historical things. I get that physical media is something that is slowly going by the wayside. But destroying literal "last copies" of stuff seems like a bad idea. And then on top of that to not even have it accessible at all to the public seems even worse. I get that they are trying to train AI. That's fine, whatever. But doing it at the cost of this kind of thing seems bad, irresponsible, potentially evil, etc.
Posted by Tree_Fall
Member since Mar 2021
1312 posts
Posted on 7/29/26 at 8:28 am to
Pretty pointless. When all is scanned and assimilated, the answer to the question "What is meaning of life" will still be 42.
first pageprev pagePage 1 of 3Next pagelast page

Back to top
logoFollow TigerDroppings for LSU Football News
Follow us on X, Facebook and Instagram to get the latest updates on LSU Football and Recruiting.

FacebookXInstagram