Podcast

Episode 91: How Books Are Destroyed

Let’s cut right to it—book people are mad at Anthropic again.

I always think it’s interesting that so many people (including the Pope) think Anthropic is “the good AI company.” Even authors, who are well aware of the whole pirating-and-scraping-books-for-training thing are more likely to say that Claude is more intelligent and personalized, less sycophantic, less generally “evil.”

But I guess when you’re in a not-being-evil contest with Open AI, it’s like you’re told to scale a brick wall then given a ladder. And that’s on top of having to pay out $1.5 billion to authors whose pirated works you used to fuel your investment-sucking machine.

So, despite the fact that the information we’re about to discuss has been out there for months and months—the most recent hubbub is over how Anthropic legally bought up millions of print books, many of them apparently rare and/or out of print, to train Claude. They partnered with a company called “Datamation” and, using a hydraulic cutting machine, chopped the spines off, scanned each page at high speed, then had them hauled off by a recycling company. Boom—donezo.

Yes, they destroyed the books. Turns out people don’t like it when you do that.

But let’s put a finer point on this, because I think what people are getting upset about here might be a little too simplistic. It’s easy to start throwing around Fahrenheit 451 and 1984, like people always do when we talk about book destruction. I have complicated feelings about that—to me it’s like Godwin’s law for books, which if you recall, is that old rhetorical thing about how you undermine your argument as soon as you bring the Nazis or Hitler into it.

But you know? A lot of us who were bringing up the Nazis and Hitler and getting dismissed for it ten years ago ended up being right. So I should grant some ground to the Bradbury and Orwell comparisons.

This situation is more than just about mulching some paper, though. It’s not just about the destruction itself. In fact, my hot take (no book-burning puns intended), is that on the driest, most practical level destroying books isn’t always a bad thing.

And even when it could be considered bad, it’s not as big of an aberration as everyone would like to think it is.

It’s the reason the books are destroyed and what (if anything) is left behind afterward that’s the issue.

And before I get back to the AI companies, I’m going to take a bit of a detour to give you a picture of the all the blasé reasons that books get destroyed.

Just your normal, everyday, ho-hum book destruction

I used to watch people have these entire meltdowns on Twitter about art projects or viral trends that involved damaging or destroying a book. There was one years and years ago where I think people were pouring milk and cereal onto books, but I could not for the life of me tell you why—if someone remembers please tell me, but I’m going to guess it was just to rage bait.

These things always throw people into enormous existential angst. And I’ll level with you—in the past I’ve been an edge-lord when it comes to the sacredness of books as material objects. I’d get really judgy when I’d see the pearl clutching, because, um like, do they know? Do they know about pulping?

In grad school, every time somebody would talk about how much they loved the smell of books and how that was why they were getting into their chosen profession, I would roll my eyes. If I were feeling really salty, I’d call them glue sniffers, because that heavenly, life-and-career-changing smell is old, dried glue and dirt.

But, plenty of people like the smell of actual dirt too. And gasoline. It takes all kinds.

Edge-lord or not—I wasn’t entirely wrong. This will come as an unpleasant surprise to some, but a sad reality for others, that the book guardians and purveyors that you love will almost all have, at one time or another, been “complicit” in destroying a book. In fact, probably many books. A huge swath of librarians, booksellers, people who work in publishing (including me), have had to destroy books at one time or another, whether by their own hand or by shipping them off to the big book farm upstate where the books can freely play and frolic.

It helps to remember context. Think of old textbooks, years out of date and full of inaccurate information. Short of holding onto a few copies for the sake of posterity, do you need to keep every single one of those?

Then there are bookworms. No, not you, oh person who loves crawling between the pages and rolling in text. I’m talking about real-life Very Hungry Caterpillars who literally eat tunnels through paper. Books can also get lice and bedbugs and create infestations in libraries, used bookstores, and your home.

And of course, in the wrong contexts, a room full of unused paper is one of the biggest fire hazards there is. Yes, I learned that one first from Ray Bradbury, but also from 9/11.

And then there’s pulping. Imagine all the books that get printed and then never get sold, isn’t picked up at any point during its life-cycle. Maybe it got remaindered first: which is when the books are sold off in bulk and heavily discounted. If you’ve ever bought a book with one of those little Sharpie hash-marks on the outside, that’s a remaindered book (and I hope you didn’t pay too much for it.) If they still don’t make it into a customer’s hands, they get sent back to the publisher or straight to a recycling center.

Then they’re pulped, which is exactly what it sounds like. Turned to pulp. Used for toilet paper. Literally.

What that says to me, before censorship or conspiracy, is that there’s is an overproduction and overconsumption issue with publishing. It makes sense. Generally, trying to figure out how many books to print in a single run is a guessing game. How many people are going to want this book? Gee, I don’t know. I hope it’s a lot!

That kind of optimism isn’t just how profit-and-loss-statements get fudged. It’s not just how author’s get their hopes dashed. It’s also how hundreds of thousands of unread books get sidelined and eventually trashed every year.

According to the Chicago Review of Books, more than 160,000 truckloads of unread books are wasted every year. That means the destruction of 10 million trees, and this was back in 2023, before the explosion of AI-generated books, and not taking into account how the number of books that real non-AI people have been producing year after year.

It’s a lot, and the a lot part isn’t just an issue for money. It’s an issue for planet Earth. Right now, the Spokane fires are making my air outside unbreathable, and so I’m feeling a little passionate about waste.

We’re so desperate for people to read books that I worry we’ve conflated a room full of books with some kind of pure morality. But if you’ve got an extra few thousand books that nobody’s going to read, the real tragedy isn’t destruction, it’s the fact that all this waste produced in the first place.

By the way—I did an interview a few years ago with Rachel Done where we got into some of the wild environmental impacts of book publishing. It’s episode 67, and you should check it out. Also, if anyone is researching sustainability in publishing and would like to chat with me about your research, message me. I’d especially like to hear what you’ve got to say about ebooks and wasted energy, because I would like a few more things to be upset about please and thank you.

As far as I’m concerned, the real tragedy of book destruction and waste doesn’t lie in aesthetics—it isn’t “ooooh I’m sending a bunch of books that have been languishing in a warehouse for years to get recycled…that’s basically the same as being a foot-soldier for the authoritarian state.”

It’s the environmental damage for one thing.

But the other thing, and the thing more relevant to this Anthropic story, is that we don’t know what we’re missing, and whatever it is, access has been taken away from us. And that’s true whether or not the books themselves get destroyed. But it is especially true when they get destroyed.

You’re not missing the point. You’re just not taking it far enough. If your first impulse is to see a book being destroyed and to get deeply viscerally upset inside of you, that is not bad. But that’s only the first level—it doesn’t go far enough. You need to get madder.

It’s not all about losing sacred objects. It’s about losing knowledge and our ability to access that knowledge and share it with one another. It’s about the loss of trust in our ability to learn and know things at all.

So why did Anthropic need all those books, anyway?

Project Panama—which is what Anthropic is calling this whole book-scanning-and-destruction project—was already reported on months ago, because information about it was included in the Bartz vs. Anthropic filing. The Washington Post wrote about it back in February, but the reason it came up again recently is because of a 404 Media report about the company ISBNdb—a book database that touts itself has having more than 111 million books, 70 percent of which are “unique” i.e., rare, i.e., not found through other means or in other databases.

Generally, ISBNdb provides data about books—it’s like a card catalog that shows people where to source them. However, the 404 post brought up some marketing copy in this report that suggested ISBNdb was branching out a little bit. Specifically, to selling bulk orders of books published before 2022 to AI companies for training.

ISBNdb is not the only company who has sold books to AI companies; in particular, Anthropic is reported to have sourced their millions through a company called Better World Books.

But ISBNdb put themselves out there and crossed their fingers the windfall would come. This post that they made, which has been taken down since the backlash, laid out the sort of system and attitudes that companies that broker these deals had. Their chipper marketing copy advertised exactly what an AI company would be looking for as they took on an operation like this.

There’s something called “model collapse” that happens when LLMs train on themselves. It’s the Multiplicity problem, the clone problem, the copy of a copy of a copy until all sense of comprehensibility is lost. The DNA breaks down, the snake eats its tail. This is a thing I’ve complained about in the past, and clearly it’s something that they are well aware is a risk.

They need real people’s real words to make any of this work. Anthropic’s internal documents said they needed these pre-LLM books, because they would teach the bots “how to write well” instead of mimicking “low quality internet speak.” Slop does not live by slop alone.

The copy that ISBNdb had posted (and since taken down) said:

“Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage […] “Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools.”

And also…

“Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.”

You can especially trust them as an AI company based on how they signal with that rule-of-threes linguistic tell. Quick, clean, classy. And according to 404 Media, they advertised orders of between one thousand and up to one million books. Notably, they also advertised that they’d have Strict NDAs for every sale, because, and I quote “The optics problem is real. ‘AI company destroys two million books’ is not a headline that generates sympathy.”

WELL YA GOT THAT ONE RIGHT, BUSTER.

They included some persuasive text to assure potential buyers that this wasn’t the same as censorship because these books were “at the end of their life-cycle” already. And, if you listened to what I said earlier about pulping, that isn’t strictly untrue. But clearly there are other concerns here. And if we take into account the feelings human beings have when we learn about books being destroyed, the optics problem is even less untrue.

Since the 404 article came out, ISBNdb has taken down that post from, and made a short post on their news page about it.

Jul-30-2026 – A note from ISBNdb:

We’ve seen the recent coverage about a marketing landing page on our site, and we understand the concern it raised. The facts: ISBNdb has never purchased, scanned, or sold a book – for AI training or anything else. We don’t train AI models, and we never have. The page was a test of market interest; no such service was ever brought to life. We’ve taken the page down. Our job is helping people find books. For more than two decades, ISBNdb has been the card catalog of the book world – the data behind how bookstores, libraries, and reading apps connect readers with titles. Data about books, not the books themselves. That hasn’t changed.

Now, do I think Anthropic had some dastardly agenda take away the world’s knowledge? I mean, I don’t think they think that’s what they’re doing. I think they’re just racing against every other company out there. They wanted to do this fast, and they didn’t want to pay for warehousing of all these millions of books. But they definitely knew people weren’t going to like it. And we know this for a fact, because they were purposely trying to keep the whole thing secret.

One internal document begged employees not to talk about the project outside of work and said: “Project Panama is our effort to destructively scan all the books in the world,” … “We don’t want it to be known that we are working on this.”

So, why was this allowed?

According to court records, as Anthropic purchased millions of print copies of books—to “build a research library”, it destroyed each print copy and replaced it with a digital copy. Their case claimed that they purchased all the printed copies “fair and square,” and that because of the fact they destroyed the books and didn’t distribute copies of them, the work was “transformative” and therefore, not a case of copyright infringement.

The judge ruled: “The print original was destroyed, one replaced the other, and there’s no evidence that the new digital copy was shown, shared, or sold outside the company.”

I’ve probably said this before, but remember: Anthropic came out on top when it came to the copyright lawsuit. The authors who are getting settlements from this aren’t getting them because of copyright infringement. They’re getting settlements because their books were taken from databases of pirated books.

This qualified as a transformative work, so according to Section 107 of the 1974 Copyright Act, if there is a transformative activity taking place. There’s a very nebulous understanding of what “transformative” means, and taking a risk has gotten a lot of especially smaller creatives into trouble. Different judges all have different opinion on what a transformative works would be—and certainly they’re not things you can profit from.

In this case, destroying the original print book and turning it into a digital copy is transformative enough that these companies are fine. The information they churn out based on this training is transformative too. This is all legal.

The funny party is that the destruction isn’t necessarily mandatory. There is technically a non-destructive way to scan all these books. It just takes a lot longer, and you have to spring for storage, because as you could probably deduce, it’s not like they can sell them off afterward. So, really, if you think about it…the destruction is just sort of incidental. Just the cost of doing business, right?

Don’t look at us, look at them

Interestingly, Anthropic hired Tom Turvey, who worked on the Google Books, to work on Project Panama. Presumably that’s not only because he knows how the actual operations of book digitization works, including non-destructive digitization, but he also knew a little bit about what they could get away with. Google ended up having to do some hoop-jumping years ago when it came to making their books projects “transformative”—i.e., snippets and excerpts rather than entire books posted online.

If you look at how the whole case shook out, destroying the books could be considered an extra layer of insurance.

But also interestingly, Google is still having copyright-related problems of their own. Google is being sued by Hachette, Cenage Group, and Elsevier, which…enemy of my enemy is my friend I guess. (For more discussion about Elsevier, go way back in HPS’s catalog to check out the episode 17 with Jahed Momand.)

A group of authors also sued Meta for the same thing last summer and lost. All of these companies are not getting pinged for copyright because, as long as they buy the books, converting them from print to digital is considered “transformative.” It’s training of the AI models is considered “transformative” as well.

I cannot stress enough how little copyright and IP is helping authors with these issues right now. And leaning on it has only given these companies a very handy way to close ranks and double down.

But Anthropic has also given all its competitors, all of whom are getting varying levels of public hate and bad press right now, an opportunity to point away from whatever the hell they happen to be doing. Elon Musk leaped out and made a big show of saying how Grok wouldn’t be like those Anthropic monsters and destroy rare books.

“I’ve asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning”

But remember: SpaceXAI is still buying up all the books and training them, though. And in order to stay in line with copyright law, they can’t distribute any of the information that was in those books outside the company. So the books belong to them, and depending on how rare they are, we don’t know if anyone else has copies. We have no way to verify that the sources the LLMs have been trained on are trustworthy, and that, when the LLMs are prompted, they’ll actually reply with accurate information.

I saw a few “pics or it didn’t happen” articles and tweets about how rare these books actually are from some of these tech bros, but I’m going to trust the smaller booksellers who have been fulfilling orders from bigger suppliers like ISBNdb and Better World Books. And when they say that one-of-a-kind, antique, irreplaceable books have all been included in these bulk sales, I’m gonna I’m gonna go with them over the crypto bros.

There is one key difference between companies hoarding the books and destroying them, though. There is one possible world where we get those hoarded books back and they somehow are accessible by the public again. In that possible world, someone who actually takes the time to look carefully finds the unique annotations or other elements that make these books special. Maybe they preserve them in an archive or digitize them for a library—like the Internet Archive, who has gotten no end of shit for (non-destructively) digitizing books for no profit.

But there’s no hope for all those books that have been destroyed, and we have no clue what we’ve actually lost culturally. And that’s the finer point on why this form of book destruction is unique and more upsetting than other kinds. Not just a loss of paper and ink.

Links for Episode 91

Want to get started? Book a 15-minute chat with Emily!