
One of my favourite science fiction novels is ‘Rainbows End‘ by Vernor Vinge. The year is 2025 (yes, seriously), and the world is a mixture of the familiar and the bizarre. Advances in medicine allow older people to become mobile and youthful again, but these newly young old people are unprepared for the world in which they awaken, full of digital wonders and wearable tech, particularly augmented reality, which is everywhere. The novel is a fascinating study of the clash between the old and the new. We see this world through the eyes of Robert Gu, a 75-year-old who has awakened cured of advanced Alzheimer’s disease, and who has to go back to school to learn how to live in his strange new life.
My favourite scene in the book involves a mass digitisation project in which physical books are not scanned page by page, they are actually fed into an industrial shredder that reduces them to tiny confetti-like pieces. Those fragments are then blown through an air stream while multiple cameras photograph every scrap; these are processed through software that reconstructs the pages from these images, effectively solving a gigantic jigsaw puzzle. If several copies of the same edition exist, the software can use overlapping fragments from different copies to reconstruct missing or damaged text accurately; the process is repeated until a full-version of the book is fully digitised. The scene is used to present the clash of the old and the new, Gu is absolutely appalled by the notion of books being destroyed, while libraries are doing this for pragmatic reasons, in a digital world the cost of housing millions of books has become needlessly cumbersome. Some unique copies are kept, but the collective corpus of human knowledge is being preserved digitally for the good of humanity. The book doesn’t particularly state that this is wrong, it just happens, although there are protesters trying to sabotage the destruction of books by pouring glue on the shredding machines, reminiscent of Luddite protests.
If you have been on social media recently, you will know why I am bringing this scene up, there have been reports that AI companies are buying millions of books, scanning them to train their models, and destroying them. The level of the response has been incredible, I don’t think that I have seen so many angry reactions in my years in the AI copyright beat. People love books, at least the people who inhabit the social media circles I frequent, so this news have hit people in a visceral way that other stories have not. The evil machines are destroying the books!
What is going on? Is it legal? What does copyright law have to do with it? I’ll try to answer these questions.
Book destruction
The reality is that this story is not new if you’ve been following the AI copyright cases closely, and let’s be honest, with over 120 ongoing cases, not many people have. The origin of the story can be found in one of Judge Alsup’s most important decisions in the case Bartz v Anthropic, which was eventually settled. As part of his ruling on whether AI training is fair use, evidence was presented that Anthropic had been buying physical books, scanning them, including the resulting digital copies into text datasets used to train their model, and then destroying them.
While Anthropic had been in conversations with publishers, they decided to go in a different route, and hired Tom Turvey, who had been an important figure in Google’s own book-scanning program (of Google Books fame). Anthropic would at some point use pirated books, but later became concerned “for legal reasons”, and so decided on a third option. According to the June 2025 order:
“Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books). Anthropic created its own catalog of bibliographic metadata for the books it was acquiring. It acquired copies of millions of books, including of all works at issue for all Authors.”
The copies were placed into a research database, cleaned of superfluous information, tokenized, and copies were kept internally as a research library. Judge Alsup ruled that this was in fact fair use, particularly the fact that the physical books had been destroyed. He comments:
“”[…] every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. […] As a result, Anthropic’s format-change from print library copies to digital library copies was transformative under fair use factor one. Anthropic was entitled to retain a copy of these works in a print format. It retained them instead in a digital format, easing storage and searchability. And, the further copies made therefrom for purposes of training LLMs were themselves transformative for that further reason, as above.”
I’ll discuss some of the legal issues here a bit later, but it’s interesting to point out the glaring issue here, AI companies have been incentivised to destroy any physical books they purchase after making a digital copy, it’s right there in the fair use ruling. Perverse incentive perhaps, but an incentive nonetheless.
You will notice that the date of this ruling is from June 2025, over a year ago, so why is this story exploding right now? At the time the fair use order was big news for many reasons, and while I do recall seeing some complaints about the book destruction, it did not reach the mainstream like it has now. The source of the current outrage is an article by 404 Media which describes all of the above, but adds new reporting stating that booksellers have been noticing an uptick in bulk book orders from unknown sources. I was actually surprised how little new detail is presented (this is not a criticism to the reporting), and why it has had such an impact. Most of the information is not new, as the book destruction was already known from last year, and the article goes on to describe that quite accurately. But I think that this passage seems to be the one that is getting more traction:
“This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.
The seller told me that, normally, on a good week, he’d sell about 20 books. Since April, he has regularly sold hundreds of books a week. While the seller didn’t have clear evidence that the purchases were being made by AI companies, the purchases made him suspect that they were.”
That’s a big if, reproduced around the Web as a fact. I have read the article several times to try to find the evidence that has prompted the current outrage, but I could not find it. The story is that booksellers are noticing an increase in sales for books from various sources, and while there is no evidence that these are going to AI companies, sellers strongly suspect it. There is citation to an article from ISBNdb which talks about book-buying as a winning AI strategy, but that article has been taken down and appears to have been written entirely by AI. The 404 article then goes on to describe Anthropic’s book destruction as a cause for the current increase in book sales. I do not think this is intentional, but if people were not paying attention to the case last year, this will come as new information, prompting the shock that ensued. (Edit to add a link to this article about the ensuing online outrage).
And I can totally see why the reaction has been so strong. I love books, I’ve written a couple, and I am the proud owner of over 800 physical books, some dating back to my teenage years, even if they’re crumbling away. Even after the Kindle and ebook revolution, I will still buy physical copies of books whenever possible, and you are more likely to find me with my head buried in some book at any given time. There has never been a time in my life when I was not reading a book. I would spend hours of my teenage years in libraries reading Clarke, Asimov, and Heinlein, the Holy Trinity of my formative years. Have I ever mentioned that I was a nerdy kid? But I digress… So the almost universal outrage has been understandable. Book burning is seen as the ultimate evil. This is what dystopian villains do in Fahrenheit 451. This is burning of the Library of Alexandria level of evil. This is what Nazis did, and you would never be on the side of the Nazis, would you?
No, I don’t think book burning is acceptable, but I am going to have to be brutally honest and admit that this story did not anger me. Part of it is that this is not new; I read Judge Alsup’s opinion when it came out. And sure, I am more on the pro-AI side of the spectrum, but there is no way that I would ever condone the wanton destruction of books, particularly if these were rare and unique books. But I have not seen evidence that this is what is actually what is happening. The 404 article states that some of the books being sold are rare, but we do not know if these rare books are being bought by AI companies and then destroyed, and this is the disconnect at the heart of the current controversy that I cannot bridge. There are a few steps here that are missing between unknown entities purchasing some rare books, and AI companies destroying them. It could be, but I am not certain enough to get outraged. I am just built in a way that is immediately suspicious of anger spirals and mob behaviour.
Books are lovely and cherished, but book destruction is not even rare; millions of unsold and unwanted books are pulped every year. Housing books is expensive, and caring for rare copies is an important part of the role of libraries. Digitisation does not have to be destructive, which is the other element of this story that has caused anger. But we do not have any evidence that unique and rare books are being destroyed in these orders, as we’ll see, depending on where the scanning takes place, they may not even need to be destroyed.
Legal issues
During the current online outrage, I have seen calls to jail AI developers (and worse), and even prompts to make this illegal, if it isn’t already. Stop the book burning!
The reality is that what is going on is part of a very important principle under copyright law called the “first sale doctrine”, or exhaustion outside of the US. This is an important legal principle which states that once a copyright owner has lawfully sold or otherwise placed a copy of a work on the market, their right to control the further distribution of that particular copy is exhausted. This means that the purchaser may resell, lend, gift, or otherwise dispose of the physical copy without the copyright owner’s permission, although copyright continues to restrict the making of new copies or other exclusive acts such as public performance or adaptation. This is a terribly important right as it allows legitimate purchasers of copyright works to resell their works, or do with them as they please.
Needless to say, without first sale and exhaustion we would not have second-hand book markets, charity shops, book-crossing, guerilla libraries, garage sales, and even giving books to your friends after you have read them. The principle is that you can do with your physical books as you see fit, that includes selling them, neglecting them, reading them until they fall apart, or letting them accumulate mould in a dusty corner of your bookcase. However, this is not an unlimited liberty, the book still has copyright, so you cannot make copies of the book to distribute them, as these are exclusive rights of the author.
As the Bartz case illustrates, buying a physical book, scanning it, and destroying the original can be fair use under US copyright law, at least in the context of mass digitisation for the purpose of training an AI model. Another mistake I have seen in the coverage is that Bartz requires AI companies to destroy the books, which is not entirely accurate, the logic here is one of substitution. Destroying the print copy means no new copy has been created, which keeps the logic of first sale intact. This gets close to a format-shifting rationale, although it is worth noting that US law contains no general format-shifting exception, and the idea that you may freely digitise an LP you own rests on surprisingly little authority.
Unsurprisingly, if this happened in other countries the results would be different. In the UK we have exhaustion, but there is not private copying exception, so creating a digital copy of a work that you own is copyright infringement. This makes all of us who at one time turned our legitimate music collection into MP3s for private consumption copyright infringers, but I digress again… In Europe, I think that these copying from physical books would be allowed as both a private copy and under the text and data mining exceptions under the DSM Directive, but I don’t think the case would be as clear-cut.
From a legal perspective there are two important takeaways. Firstly, there is indeed some sort of legal incentive to destroy books, which I think should not exist. Secondly, just because we dislike the current state of affairs we should not get rid of exhaustion and the first sale doctrine, the benefits to the public to lose just because of AI companies behaving badly.
Edit to add: Someone posted on X that the demand may also be coming from people who assume AI labs will be purchasing physical books in the future, and this makes a lot of sense. Old books have become a valuable commodity, good time for book sellers.
Concluding
Not for the first time when dealing with AI, I must confess that I’m a bit apprehensive about writing and publishing this article given the strength of sentiment and the level of public outcry that the news has generated. I can just imagine a badly-intentioned person reading the above and claiming “Andres is in favour of burning books!” Before you reach for your pitchforks, that is not what I’m saying. If evidence comes to light that AI companies are destroying unique books, you can count me among the first to publicly state my complete disgust, and I’ll join the mob trying to pour glue on the scanners.
But I just haven’t seen evidence that this is happening as described. The story has also been mutating into some sort of weird urban legend that has no bearing on reality. The story has a legitimate origin in the Bartz v Anthropic case, but it has changed into the claim that AI companies are destroying books so that they can “erase history” and control what people believe. This strange version has no substance. It is a fantasy world created to anger people and get more clicks; this is engagement farming gold.
The reality is probably less outrageous than the online mob implies. There is no doubt that some destructive scanning took place, but there is no need to continue doing that, particularly if the digital copies of those books are deemed to have been made legitimately. It is likely that AI companies are indeed buying books in bulk, but we don’t know where the scanning is taking place, and most importantly, whether it is destructive or not. As I mentioned, I think that this type of non-destructive scanning would fall under the TDM exception in the EU, and probably in many other countries.
I am left thinking about the destructive scanning in ‘Rainbows End’. I think that Vinge knew exactly what a gut punch that scene would be; I still think about it often. Vinge doesn’t pass judgement, which is just another level of why that novel is so good. It is also quite prescient in some ways that we may find uncomfortable beyond the AI scanning element. Reading is in decline around the world, even as more people than ever buy ebooks and listen to audiobooks, and there is a new generation bypassing books entirely in favour of video. I don’t like that future, heck, I think TikTok is an abomination that is destroying people’s attention span. In the end we are indeed living through this clash of the old and the new, the book and the machine. Who will win?
I don’t know who will, but I will keep reading nonetheless. I think that ‘Rainbows End’ is due a re-read anyway.
1 Comment
AI & Books: notes from a moral panic – Peter Gasston · July 31, 2026 at 3:44 pm
[…] On July 21st, 404 media published AI Companies Are Buying Tons of Old Books Because They’re Free of AI Slop. It describes the practice where, to train their models with ‘premium’ human-written data, AI companies are using brokers to buy copies of books they don’t already have in their datasets, which they scan using a destructive method (they’re not allowed to keep a physical copy after scanning due to copyright law). […]