When AI 'Eats' Out-of-Print Books: The Physical Cost of LLM Data Hunger
Australian booksellers are protesting the destructive scanning of rare books to train AI. This article explores how large language models acquire data, the depletion of high-quality text, and the ethical trade-offs between digital acceleration and physical preservation.

Recently, antiquarian booksellers in Australia have raised an alarming protest. According to reports, rare books and out-of-print documents are being roughly dismantled, scanned, and even subjected to "horrific destruction" to acquire data for training artificial intelligence (AI).
This might sound surreal, but it is a real and growing pain point in the tech world today. Let's explore how AI's "appetite" is impacting the physical world.
Dismantling Out-of-Print Books: What Exactly is AI Eating?
The Large Language Models (LLMs)—the AI systems you chat with fluently—derive their core capabilities from "reading" massive amounts of text. In the past, they primarily consumed free data from the internet, such as Wikipedia, news articles, and forum posts.
However, as AI becomes more advanced, the "high-quality corpora" on the internet—text that is logically rigorous and free of fluff—are running out. Consequently, data mining companies have turned their attention to the physical world: physical books, rare documents, and out-of-print journals that have not yet been digitized.
Imagine this scenario: an antiquarian bookstore owner is approached by an AI data buyer. The buyer says, "I'll pay top dollar for this out-of-print 19th-century botanical guide, but we need to cut off the spine and put it through a high-speed scanner." The owner refuses heartbrokenly, "The value of this book lies not just in the words, but in the two-hundred-year-old binding and paper. Cutting it will destroy it."
This is the reality Australian booksellers are currently facing. In the pursuit of scanning speed, precious bindings are destroyed, and books are reduced to disposable data consumables.
Simply put, AI's "intelligence" is not generated out of thin air; it is fed by consuming the crystallized knowledge of humanity. When the free lunch of the digital world ends, it begins to reversely extract the cultural heritage of the physical world.

Why Focus on Physical Books?
Some might ask: Can't AI generate its own data to train the next generation of AI?
Current technology shows that training next-generation AI with AI-generated data can easily lead to "Model Collapse." It's like making a photocopy of a photocopy; the image and logic become increasingly blurred. Therefore, original, time-tested, high-quality human text remains irreplaceable "golden data."
This means that LLM training is facing a crisis of "high-quality data depletion." Those physical books sitting deep in libraries or in the hands of private collectors have become the last gold mine. The physical destruction protested by booksellers is, in essence, the predatory extraction of physical cultural assets by digital capital.
The Trade-off Between Digital Frenzy and Physical Costs
In my view, this is not just a simple copyright dispute, but an ethical trade-off between "digital immortality" and "physical preservation." As booksellers emphasize, books are more than just objects carrying information; the paper, marginalia, and binding craftsmanship themselves are part of history. Dismantling a book just to extract the text is akin to destroying the historical artifact for the mere information, erasing the physical attributes of the cultural relic.
Looking at it from another angle, this brings to mind the medieval practice of creating a "palimpsest." Because parchment was so expensive, monks would scrape off ancient Greek classics to copy religious texts. Today, rare books are physically dismantled to train AI. Is this not a modern version of "scraping"? In our pursuit of future intelligence, we are erasing the physical traces of the past.

How Should the General Public View and Respond?
This issue may seem distant, but it is closely related to us all. If rare books are destroyed on a large scale, the cost for ordinary people to find well-preserved out-of-print books in the second-hand market will become extremely high in the future. At the same time, it reminds us that the AI products we use daily may hide unknown cultural and ecological costs behind them.
It is worth being vigilant about the extreme techno-optimism that suggests "everything can be sacrificed for AI development." Technological progress should not come at the expense of destroying human cultural heritage. When paying attention to AI stocks or related investments, please also be mindful of data compliance and supply chain risks (Note: Related investment analysis is for reference only and does not constitute professional advice).
Currently, regions like the European Union have begun implementing stricter AI regulations, such as the EU AI Act. How to balance data acquisition and physical preservation in the future remains a question for global legislators to answer.
One-sentence summary to share: AI training is facing a depletion of high-quality data, leading to the destructive scanning of rare physical books; digital acceleration should not come at the cost of destroying physical cultural heritage.
Discussion: If you had a precious, out-of-print old book, would you be willing to let its spine be cut for destructive scanning in exchange for its "digital immortality"? Welcome to share your thoughts in the comments.