乐于分享
好东西不私藏

为训练AI软件,Anthropic 不惜销毁数百万册书籍,并且不想让外界知晓……

为训练AI软件,Anthropic 不惜销毁数百万册书籍,并且不想让外界知晓……

研发机器学习与生成式软件的企业对书籍的掠夺,并不仅仅停留在比喻层面。至少有一家公司,真的在粉碎数百万本实体书籍,用来投喂自家聊天机器人。上月《华盛顿邮报》曝光,AI 巨头 Anthropic 开展了一项名为 “巴拿马计划” 的大型项目:投入数千万美元大量收购旧书,裁切、扫描之后将纸张制成纸浆,从书籍中提取的扫描数据被用于训练模型。

Anthropic 此前已经因为盗版数百万本电子书陷入舆论漩涡。但在该案中法官模棱两可的判决,被 Anthropic 律师视作可利用的法律漏洞。法官裁定:如果利用书籍训练 AI 属于 “转换性使用,就在法律允许范围内。这类似于用书教育孩子,也对应一条规则:一旦买下实体书,你便可以随意处置它 —— 二手书店能够存在,正是依托这项法律先例。

巴拿马计划” 充分利用了这个漏洞。Anthropic 花费巨资,从图书馆、线上二手平台、The Strand 这类实体旧书店大批量收书,搭建巨型藏书库。《华盛顿邮报》报道配图里,可以看到堆满书籍的大型仓库。Anthropic 还聘请了一位前谷歌图书项目员工主管这项工作。

整套流程如下:数百万本收购而来的书籍被送上工业流水线。机器切掉书脊,将装订成册的书本拆成散页。高速扫描仪采集每一页文字。扫描完成后,实体书籍就被送去制浆回收。

在持续进行的版权诉讼中解封的一份内部规划文件直白写明项目目标:巴拿马计划,是我们对全世界所有图书开展破坏性扫描的行动。文件同时警示员工:我们不希望外界知晓我们正在推进这件事。

这项计划和上一轮大规模图书扫描项目 —— 谷歌图书,存在根本性区别。谷歌承诺向公众开放扫描文本。而巴拿马计划产出的扫描文件,只保存在 Anthropic 私有服务器内。读者、图书馆、作者都无法访问。

不妨仔细思考这件事:数百万本实体书籍,其中一部分珍稀、绝版、无可替代,就此永久消失。它们仅存的形态被锁在企业私有数据库中,只为驱动商业化聊天机器人。这些文本并非面向公众存档,唯一目的是让 Anthropic 的软件产生更多收益。

几年前,二手书商就开始留意到奇怪的大批量订单。书商向《华盛顿邮报》透露,他们收到一次性采购数万册图书的订单,往往附带严格保密协议。中介走遍图书馆旧书售卖场、旧货店、图书批发市场,满足 Anthropic 采购需求。

图书行业很多从业者心生疑虑:珍本、学术专著、冷门诗集、绝版小说 —— 这类几乎不会大批量流通的图书,也被一扫而空。

Anthropic 赖以支撑的法律逻辑,依托美国版权法中的首次销售原则:一旦合法购买书籍实体本,你拥有这件实物,有权销毁它。叠加法官认定 “AI 训练属于转换性合理使用,企业认定整套破坏性扫描流程受法律保护。

但合理使用原则,从来不是为了允许私有企业系统性销毁文化作品实体,用来打造独占的专有数据集。图书馆保存书籍,档案馆保存书籍;企业没有义务传承文化遗产。

谷歌图书上线时,引发全球激烈争论,但它的核心愿景是打造面向全人类开放的通用数字图书馆。巴拿马计划恰恰相反:清除流通中的实体书籍,搭建封闭、专有、仅一家科技公司能够使用的数字图书馆。

试想这样一种未来:无数绝版书籍,仅仅存在于私人 AI 服务器中。一旦这家公司决定删除扫描文件,或是破产,这些作品将彻底消失。没有图书馆留存实体本,普通公众无从阅读。实体书本早已销毁,数字副本也不会对外共享。

这并非抽象的反乌托邦猜想,正是巴拿马计划想要实现的目标。

随着 AI 行业竞争白热化,其他企业争相搭建规模更大的私有文本数据集。如果这套模式被广泛效仿,这场悄无声息、工业化的纸质文献销毁行动,只会愈演愈烈。


附原文:

Anthropic didn’t want us to know that they were destroying millions of books to feed their software.

James FoltaFebruary 6, 2026

Companies making machine learning and generative software aren’t just metaphorically ripping off books. In at least one case, they’re rather literally shredding millions of physical books to feed to their chatbots. As uncovered last month by The Washington Post, AI giant Anthropic ran a massive program called Project Panama where they spent tens of millions of dollars to hoover up used books, which they then sliced, scanned, and pulped. The scanned data removed from the books was then used to train their software.

Anthropic has already been in the hot seat for getting caught pirating millions of digital copies of books. But the judge’s equivocating ruling in that piracy case created a loophole, according to Anthropic’s lawyers. If the books training AI were used in a “transformative” way, the judge ruled, it was legally aboveboard, akin to using books to teach kids or how you can do what you want with a book once you buy it—a legal precedent that allows for secondhand bookstores, for example.

Project Panama capitalized on that loophole. Anthropic spent a bundle at libraries, online secondhand stores, and used bookstores like The Strand to build out a massive library—the Post’s article includes images of huge warehouses filled with books. Anthropic then hired a former Google Books employee to oversee the operation.

Here’s how it worked: Millions of purchased books were fed onto an industrial assembly line. Machines sliced off their bindings, turning bound volumes into loose stacks of pages. High-speed scanners captured the text on every sheet. Once scanned, the physical books were sent to be pulped and recycled.

An internal planning document unsealed during ongoing copyright litigation laid out the goal plainly: “Project Panama is our effort to destructively scan all the books in the world.” The document added a warning to staff: “We don’t want it to be known that we are working on this.”

There’s a crucial difference between this project and Google Books, the last massive book-scanning effort. Google promised to provide public access to scanned texts. The scans from Project Panama exist only within Anthropic’s private servers. No readers, libraries, or authors can access them.

Think about that for a moment: Millions of physical books, some rare, out-of-print, and impossible to replace, permanently erased. Their only surviving form is locked inside a private corporate database, deployed to power a commercial chatbot. The texts are not archived for public benefit; they exist solely to make Anthropic’s software more profitable.

Secondhand booksellers began noticing strange bulk orders a few years ago. Book dealers told the Post they received orders for tens of thousands of volumes at a time, often with strict non-disclosure agreements attached. Brokers scoured library sales, thrift shops, and wholesale book markets to fill Anthropic’s requests.

Many in the book trade grew suspicious. Rare books, academic monographs, obscure poetry collections, out-of-print novels—titles that rarely sell in bulk—were swept up in the purchases.

The legal logic Anthropic relies on hinges on the first-sale doctrine in United States copyright law. Once you legally purchase a physical copy of a book, you own that physical object, and you are permitted to destroy it. Combined with the judge’s finding that AI training qualifies as transformative fair use, the company believed the entire destructive scanning pipeline would be protected.

But fair use was never intended to enable private corporations to systematically eliminate physical copies of cultural works to build exclusive proprietary datasets. Libraries preserve books; archives preserve books. Corporations do not exist to preserve cultural heritage.

When Google Books launched, it sparked fierce global debate, but the core vision was a universal digital library open to humanity. Project Panama is the inverse: removing physical books from circulation to create a closed, proprietary digital library accessible only to one tech company.

Imagine a future where countless out-of-print books exist nowhere except locked inside private AI servers. If that company decides to delete the scans, or goes bankrupt, those works vanish entirely. No library holds a physical copy. No member of the public can read them. The physical book has been destroyed, and the digital copy is not shared.

This is not abstract dystopian speculation. It is exactly what Project Panama set out to accomplish.

As AI competition heats up, other firms are racing to build larger private text datasets. If this model is widely copied, the quiet industrial destruction of printed literature will accelerate.