Anna's Archive—the shadow library project that has quietly become one of the most comprehensive collections of digitized texts online—published a blog post this month that should matter to anyone who depends on AI tools for research or knowledge work. The argument: AI companies are bulk-acquiring rare physical books as training data, in some cases using destructive scanning methods that leave the original volumes in pieces, effectively removing them from public access forever. Here's the pitfall that deserves immediate attention, before you scroll past this thinking it's an abstract preservation debate: even if you never plan to touch a shadow library, the books disappearing into private training sets today are the sources that will never be citable tomorrow—and the AI tools your team pays for are drawing on a knowledge layer that has no public audit trail.

The 833-comment Hacker News thread this generated is one of the more substantive tech discussions in recent memory, cutting across copyright law, preservation ethics, and the genuinely uncomfortable question of who owns humanity's recorded knowledge once it becomes training data.

For small teams and solo operators, this matters less as an abstract cultural crisis and more as a concrete shift in how verified, traceable knowledge gets created and gatekept going forward.

What Is This Actually?

Anna's Archive is a meta-search engine and mirror for shadow libraries—primarily Library Genesis, Sci-Hub, and Z-Library—plus its own original scanning projects. It has indexed tens of millions of books, papers, and documents that are technically under copyright but hosted outside the traditional publishing system. The project operates in a legally gray zone and is blocked or restricted in several jurisdictions, but it's one of the few organizations actively working to create a comprehensive public backup of human knowledge outside corporate or government control.

The specific concern raised in their August 2026 post is narrower and more alarming than the general shadow-library-versus-publishers debate. AI companies are now competing for physical books—specifically rare, out-of-print, or single-copy volumes—to use as training data. The mechanism at the center of this is destructive scanning.

When a high-throughput digitization team needs to process thousands of books quickly, cutting the spine and feeding loose pages through a document feeder is dramatically faster than using a flatbed or overhead camera rig. For a paperback edition of a commercially available novel, destructive scanning is annoying but not catastrophic—there are other copies in circulation. For a regional monograph from 1931 with three known surviving copies, destructive scanning by a private buyer means the world now has two. The physical damage is permanent.

The pipeline, as Anna's Archive describes it, looks roughly like this: AI companies, or third-party data brokers supplying them, acquire physical books through used book dealers, estate sales, and private auctions. Some portion of this acquisition specifically targets out-of-print and rare material, precisely because it doesn't already exist in digital form and therefore isn't subject to API access restrictions or existing licensing arrangements. Books get scanned—sometimes carefully, sometimes not—and the text enters proprietary training pipelines. The physical books, damaged or not, don't necessarily end up in libraries or archives afterward.

Anna's Archive's argument is essentially a race condition: if these books are going to be consumed into AI training data anyway, the world would be better served if they were first scanned using preservation-quality methods and added to publicly accessible archives before disappearing behind a corporate data moat.

It's worth being honest about what the post can and can't establish. The scale of AI-specific rare book destruction—how many rare volumes are genuinely being destroyed through AI data acquisition versus the much older, chronic problem of storage neglect, estate clearances, and library deaccessioning—is not something Anna's Archive can precisely quantify. Their post is directional rather than forensic: it observes acquisition patterns and makes a plausible argument about where this is heading. That doesn't make the concern less valid. It does mean holding the most alarming framing loosely while taking the structural problem seriously, which is where we land after working through the full picture.

The project is calling for volunteers, funding, and institutional cooperation to scan rare holdings—particularly those in smaller regional libraries, historical societies, and private collections—before this material either becomes unavailable or becomes exclusively accessible through AI systems that have no obligation to share it.

Why This Matters Right Now

The timing isn't coincidental. The legal environment around AI training data and digitization has deteriorated rapidly over the past two years.

In the US, the Internet Archive's controlled digital lending program was ruled not protected by fair use (the Hachette v. Internet Archive decision), forcing them to remove millions of books from their lending library. That ruling had a chilling effect on every institution that was cautiously scanning and sharing material under a legal theory that courts ultimately rejected.

Meanwhile, AI companies facing copyright lawsuits have mostly been arguing that training on copyrighted material is transformative fair use—a position courts haven't definitively settled yet, but one that has emboldened large-scale acquisition. The legal asymmetry here is striking. The Internet Archive, a non-profit trying to give people access to read books, got sued and lost. AI companies processing those same books into products valued in the tens of billions are still litigating, on better-funded legal teams, with more favorable case postures. The legal system is producing an outcome where the most commercially valuable use of books faces the weakest constraints, while the most publicly beneficial use faces the strongest. Whether that asymmetry holds up over time is genuinely uncertain.

What's changed specifically over the past 12 to 18 months is the training data focus shifting from web text to books. Early large language models were trained heavily on web crawls—Common Crawl, Reddit, news sites. That source is well-indexed and easily accessible, but it's becoming more homogeneous as the web itself increasingly consists of AI-generated content. Books represent a different signal: denser, more carefully edited, drawing on domains and time periods not well-represented on the open web.

The demand for book-sourced training data has intensified at exactly the moment when legal access to that data is becoming more restricted. Physical acquisition—buying actual books rather than licensing digital editions or scraping existing archives—is a way around those restrictions. The physical book, once purchased, is yours to do with as you choose. The copyright question is separate from the ownership question, and that gap is being exploited.

The result is that the knowledge layer underpinning the AI tools small teams use for research, summarization, and content generation is increasingly sourced from a pool of documents with no public audit trail. You can't cite the rare 1940s industry trade journal that your AI tool "knows about" because there's no way to verify it even exists in a form you can access and inspect.

Practical Implications for Small Teams

This plays out differently depending on how your team uses AI and what kind of knowledge work you do.

Research-heavy agencies and consultants are probably most exposed. If your workflow involves asking AI tools to summarize historical trends, pull context from domain-specific literature, or synthesize findings from obscure sources, you're already relying on training data you can't inspect. What this development adds is the possibility that sources your AI tools "know" from their training could become permanently inaccessible for verification—not because they were online and got taken down, but because the physical originals were processed into private training sets and the books themselves no longer exist in accessible form. When a client asks you to show your work, that creates a problem.

The practical implication is not to panic but to start building source-verification habits now. When an AI output makes a claim that would be valuable to verify or cite, develop a consistent workflow for tracing it to something a human can actually access: the Internet Archive, HathiTrust's public domain collection, a library catalog, or directly through the original author or publisher.

Content teams and solo creators face a different version of the problem. The writing assistants and research tools they use are going to diverge in capability based on what training data each one has access to. A tool with richer book-sourced training data will produce better domain depth. But that richness may have come at a cost that your subscription fee doesn't make visible. You might be benefiting from—and implicitly endorsing—a data acquisition practice your team would find uncomfortable if you knew the specifics.

Small software teams building on AI APIs have a more specific concern. If you're building a product that leans on an LLM's knowledge of domain-specific literature—legal precedents, medical research, historical technical standards—you need to understand where that knowledge comes from and what happens when its accuracy is challenged in a client relationship or legal context. A tool that "knows" something because it's in a rare scanned book is in a different evidentiary position than one drawing on peer-reviewed, openly accessible sources. Your downstream liability exposure depends on that distinction.

Teams in education, nonprofits, and civic tech face the most direct version of the equity concern running through the Anna's Archive post. If the digitization of rare knowledge migrates from public archives to private AI datasets, the asymmetry between organizations that can afford enterprise AI subscriptions and those that can't becomes a knowledge asymmetry, not just a productivity one. A public school researcher using free tools and a hedge fund researcher using frontier AI enterprise access would, in this scenario, be working from fundamentally different information bases—and only one of them would know it.

There's also a quieter implication for anyone doing competitive research or market intelligence. Historical data—old trade journals, out-of-print industry reports, regional business histories, technical manuals—is exactly the category of material being targeted for AI training. A small agency doing deep historical research today has access to physical and partially-digitized sources. That same research done in three to five years may be significantly easier if AI tools have ingested this material into something you can query, or significantly harder if the originals no longer exist in accessible form. The outcome depends entirely on whether digitization happened for the public or for a private training pipeline.

How to Respond / Act on This

The actionable response isn't complicated, but it requires deliberate choices rather than passive consumption.

Start with your research source audit. If your team regularly uses AI tools for research, spend thirty minutes mapping what kinds of sources those tools are actually good at citing. Tools that return generic web results are in a different position than those claiming domain expertise from literature. For the latter, build a parallel habit of confirming key claims through open archives. The Internet Archive, HathiTrust's public domain collection, and Project Gutenberg are all free and legally clear. Make these defaults, not fallbacks.

Contribute to or fund preservation efforts if this aligns with your mission. Anna's Archive explicitly accepts donations and volunteers. For teams that depend heavily on open knowledge access—academics, nonprofits, researchers, journalists—this is worth treating as infrastructure spending rather than charity. The Internet Archive is a 501(c)(3) in the US; donations are tax-deductible and support one of the few institutional barriers to complete privatization of the historical record. If you have access to physical book collections through a local historical society, small library, or private collector, that access is worth more than you might think to a coordinated scanning project.

If your team does any scanning of physical documents—client records, historical research material, your organization's own archives—switch to non-destructive methods and preservation-quality formats. The PDF/A format is designed for long-term archival. Tools like ABBYY FineReader (around $15/mo for a subscription, or a flat-fee desktop license) and Adobe Acrobat both support non-destructive scanning workflows with solid OCR output. Hardware like the CZUR ET series of overhead scanners—starting around $200 for entry-level models—allows book scanning without spine damage. For anything that can't be replaced, the extra setup time is worth it.

Be deliberately skeptical of AI-as-oracle. The tendency to treat LLM outputs as authoritative is strongest when the topic is obscure enough that most users can't immediately fact-check the claim. Rare historical documents, domain-specific technical knowledge, and out-of-print sources are exactly where AI hallucination combines most dangerously with inaccessible originals. If an AI confidently cites something you can't find anywhere in accessible form, that's a signal that warrants follow-up—not a confirmation that the source exists.

Monitor the legal landscape. The training data copyright cases moving through courts in the US and EU will have direct effects on what AI tools can claim their knowledge is based on. Teams building products on AI APIs should have someone tracking these developments. Not because you need a lawyer on retainer, but because the terms of service and effective knowledge base of the tools you depend on may change substantially depending on how these cases resolve. What the tool knows today may not be what it can legally claim to know in 18 months.

Knowledge Access Tools: A Comparison

For teams that need reliable research access beyond what AI tools can provide—especially for historical, academic, or domain-specific material—here's how the main options compare:

Tool Best for Free plan Starting price Key differentiator
Internet Archive Public domain books, historical web, periodicals Yes Free Largest publicly accessible digital library; legally clear for public domain
HathiTrust Digital Library Academic research, full-text search of scanned books Yes (public domain) Free Backed by major research universities; strong OCR quality
Project Gutenberg Classic literature, pre-1928 public domain texts Yes Free Oldest digital library; clean plain-text and epub formats; no DRM
Google Books Snippet search across millions of titles Yes Free Largest index of scanned books; partial access to in-copyright works
Semantic Scholar Academic papers and scientific literature Yes Free AI-powered research graph; 200M+ papers; strong citation linking
Zotero Citation management and team research organization Yes (2GB) ~$20/mo storage Best reference manager for teams; browser integration; open source
Elicit AI-assisted literature review Yes (limited) ~$10/mo Extracts structured data from papers; good for systematic research
Anna's Archive Rare books, out-of-print material Yes Free Most comprehensive shadow library index; legally gray

Our take: for most small teams, the combination of Semantic Scholar for papers and Internet Archive for historical documents covers the majority of legitimate research needs without legal risk. Zotero as your organizational layer on top is genuinely underrated for collaborative citation management—and it's open source, which matters if you care about tool longevity. What tripped us up in researching this piece is how often people jump straight to AI tools for research without first checking whether the source material actually exists in HathiTrust or the Internet Archive at no cost. The open archive ecosystem is richer than most people realize.

What the HN Community Is Saying

The 833-comment thread is long but worth parsing because it splits in interesting and revealing ways.

The strongest skeptical voice asks consistently for specifics: which AI companies, in what volume, buying from which channels? This is a fair demand. The Anna's Archive post is directionally credible but light on named actors and verifiable figures. Several HN commenters with apparent knowledge of used book markets pointed out that rare book destruction is actually a longstanding problem driven as much by estate clearances, storage cost constraints, and library deaccessioning as by any AI-specific demand. The AI angle may be drawing attention to a newer acceleration of a much older process—which matters for how you respond to it.

The preservation advocates in the thread—and there are many—largely agreed that the call to scan is correct regardless of whether the specific AI framing is precisely accurate. Books disappear from the historical record for many reasons. Accelerating high-quality digitization is defensible on its own terms, whether or not AI data brokers are the primary threat.

The most substantive sub-thread concerned the legal asymmetry. Several lawyers and law students in the discussion pointed out that the current legal framework seems to be producing an outcome where the most commercially valuable use of books—training AI models—faces the weakest legal constraints, while the most publicly beneficial use—lending them—faces the strongest. That's not a bug in how copyright law is being applied; it's arguably a feature of how corporate litigation resources shape legal outcomes over time.

The tech community's pragmatic wing offered something more actionable than the debate. Several people in the thread are organizing scanning projects and calling for anyone with access to rare collections—local historical societies, small university libraries, private collectors—to make contact with preservation groups. The distributed volunteer model is explicitly positioned as the only one that produces publicly accessible results.

What the thread doesn't produce is consensus on whether this is a crisis or a slow-moving structural problem. Our read: it's the latter in aggregate, but with specific categories of material—single-copy regional histories, pre-digital technical literature, non-English language archives—where the crisis framing is genuinely accurate and the window for action is shorter than most people assume.

Risks and Things to Watch

The non-neutral narrator problem. Anna's Archive is not a disinterested party. They are a shadow library with their own legal exposure and their own interest in framing AI companies as the villains in a narrative that positions Anna's Archive as the hero. That doesn't make their underlying concern wrong—but it means their specific claims about AI company behavior deserve skepticism proportional to their interest in making those claims. We'd be cautious about treating any single advocacy piece, even a well-argued one, as the definitive account of what's actually happening at scale.

Scanning economics are harder than they look. If this post inspires teams or individuals to embark on digitization projects, the practical economics are worth understanding before you commit. Non-destructive book scanning to archival quality is slow—roughly 10 to 30 minutes per book for a careful operator with decent equipment. Even with a good overhead scanner, digitizing a meaningful collection of rare books is a significant time investment. Organizations that want to contribute at scale should be realistic about throughput and plan accordingly, rather than starting projects that stall halfway through.

The legal gray zone compounds. Anna's Archive operating outside copyright law means that material scanned and added to their index has different downstream legal status than material added to HathiTrust or the Internet Archive. Teams that want to use scanned material in commercial workflows should be conscious of where that material originated and whether their use creates exposure. For individual research, the risk is low. For building products on top of that content, it's worth a conversation with a lawyer.

Vendor lock-in to AI knowledge. There's a subtler risk for teams deeply integrated with specific AI tools: as those tools' knowledge bases become richer through access to book-sourced training data, migrating to a different tool becomes harder. If your team's workflows assume a specific AI's domain knowledge and that knowledge is proprietary, you're exposed to both price changes and to whatever legal outcomes change that tool's capabilities. Diversifying across tools and maintaining access to primary sources is a hedge against this.

The public domain timeline is worth tracking. Works published in the US through 1927 are now firmly in the public domain. Works from the 1930s through the 1960s are in various stages of entering the public domain depending on registration status, renewal history, and jurisdiction. Organizations with interests in mid-20th century material should monitor this—material entering the public domain can be legally scanned, shared, and incorporated into open archives. It's not a fast-moving shift, but it's a consistent one that opens up new preservation opportunities each year.

Frequently Asked Questions

What exactly is "destructive scanning" and how common is it?

Destructive scanning means removing a book's binding so individual pages can be fed through a document scanner rather than held flat under an overhead camera. It's significantly faster than non-destructive methods and produces cleaner scans with fewer shadows or distortions. The technique is common in commercial digitization workflows—notably used in portions of Google's early book scanning project and widely used by document management companies processing business records. For common books with many surviving copies, it's a pragmatic trade-off. For rare books with few surviving physical copies, it's genuinely irreversible. How common it is specifically in AI training data acquisition is not publicly documented; that's a real gap in the public record that makes it hard to assess the scale of the specific concern Anna's Archive is raising.

Is Anna's Archive legal to use?

The short answer is that it depends on your jurisdiction and use case. Anna's Archive aggregates material from Sci-Hub, Library Genesis, and Z-Library, none of which operate with the consent of publishers or rightsholders. Downloading copyrighted books from these platforms is technically copyright infringement in most jurisdictions. Enforcement against individual end users has been rare in practice, but it does happen. For teams and businesses, the risk is higher than for individuals—particularly in industries where IP compliance is audited or expected. For genuinely out-of-print and effectively inaccessible material, the moral calculus is different from the legal one, but those remain distinct questions.

Are AI companies required to disclose what books they train on?

Not currently in most jurisdictions, though this is changing. The EU AI Act includes provisions that would require disclosure of training data categories for high-capability models. In the US, no such requirement exists yet, and the legislative environment is uncertain. Several ongoing copyright lawsuits have produced discovery requests that are beginning to surface some specifics about training corpora, but the overall picture remains substantially opaque. This opacity is part of what makes the Anna's Archive concern difficult to verify or dismiss—the lack of transparency is both a practical problem and a symptom of the regulatory gap.

If AI companies have already scanned rare books for training data, can those scans be accessed publicly?

No—and that's the core of the problem. Training data isn't a searchable archive; it's material processed and transformed into model weights. You can't retrieve the original scan of a rare book from a trained AI model. The knowledge exists in a distributed, abstracted form, but the source document doesn't become accessible to others as a result of being in the training set. The book has been effectively consumed rather than preserved. This is categorically different from what Google Books did (where scanned books remained searchable and partially viewable) or what the Internet Archive does (where scans are available for borrowing). An AI that "knows" something from a rare scanned book provides no path back to the source.

What's the best way for a small team to access rare or out-of-print research material legally?

Start with HathiTrust, which has deep collections of out-of-print material from research libraries—much of it in the public domain and fully accessible, with in-copyright material available for full-text searching but not downloading. The Internet Archive's book lending program has been curtailed by the Hachette ruling but still offers access to many titles. For material that doesn't appear in either, Interlibrary Loan through a local public or academic library remains remarkably underrated—rare books can often be borrowed or scanned on request by a librarian, at no cost to the patron, even from collections far from your location. Most people don't know this service still works well for genuinely obscure material.

Does this affect which AI tools are actually better for research?

Indirectly, yes. Tools trained on richer book-sourced datasets tend to have better domain depth in areas not well-covered by the open web—specialized history, pre-digital technical literature, regional and non-English sources. But "better at research" and "more opaque about sourcing" often come together in AI tools with privileged training data access. A tool that knows more about an obscure topic because it trained on a private scan of a rare book is harder to fact-check than one drawing on accessible sources. The capability gain comes with a verification cost that matters for professional research where claims need to be supported by something a client or peer can inspect.

Should small businesses be worried about copyright liability when using AI tools trained on books?

Current legal consensus is that end users of AI tools don't inherit copyright liability from training data. The liability questions are between AI companies and rightsholders, not between AI companies and their customers. This could change depending on how ongoing litigation resolves, but for now, using an AI tool for your business is not the equivalent of downloading a pirated book. The separate question is whether AI-generated content that closely reproduces specific copyrighted works creates liability—that's a different analysis that applies regardless of how the content was generated, and it's the one more likely to affect small teams directly.

What happens if AI companies win the argument that training on books is fair use?

Paradoxically, it might not solve the preservation problem—and could accelerate it. If training on books is definitively ruled fair use, it reduces the legal friction on bulk acquisition and scanning, which could increase the volume of rare books being processed into private datasets. Whether those datasets are made publicly accessible is entirely separate from the legal question of whether creating them was permissible. An AI company could win the fair use argument, continue scanning rare books at scale, keep all resulting scans proprietary, and face no obligation to share them with libraries or researchers. That outcome would validate the AI companies' legal position while leaving the preservation problem fully intact.

Final Verdict

This story is about more than books. It's about who controls access to the knowledge layer that AI tools are built on—and whether the digitization of human knowledge is happening in a way that serves the public or primarily serves the companies doing the digitizing.

For small teams and solo operators, the immediate practical impact is limited. The AI tools you use tomorrow won't suddenly break because a warehouse of rare books was scanned last month. But the directional trend is real and worth watching, because the gap between what AI systems have ingested and what humans can independently access is widening. That gap matters most exactly when you need to verify, cite, or build on AI output—which is precisely the use case that creates professional value and professional liability.

Recommendations differ by team type. Research-heavy agencies and consultants should act now: audit your AI research workflow, build source-verification habits into your process, and diversify toward open archives. The Internet Archive and HathiTrust cost nothing and provide a legally clear fallback for historical material. Teams for whom research is occasional rather than core can watch and wait—this won't bite you acutely in the next six months. Teams building products on AI APIs should add training data provenance to their vendor evaluation criteria; not every provider is equally opaque, and disclosure requirements from the EU AI Act will create more differentiation here over the next 12 to 24 months.

For those inclined to contribute to the preservation side: Anna's Archive takes donations and volunteers, the Internet Archive is a US non-profit accepting tax-deductible contributions, and local historical societies with physical book collections are often looking for scanning support without knowing where to find it. The volunteer scanning model is slow and imperfect. It's also the only model that reliably produces publicly accessible results rather than feeding a private training pipeline.

What this signals more broadly is that the AI training data acquisition push is entering a phase focused on physical and historical material—harder to find, legally harder to contest, and not available through the web crawl methods that worked for the previous generation of model training. The organizations paying attention to this now, and building workflows that don't depend entirely on opaque AI knowledge claims, will be better positioned when the legal and regulatory environment eventually catches up to where the data acquisition actually is.