AI labs are bulk-buying second-hand books, training models on the contents, then destroying them. Booksellers fear rare and out-of-print titles are being pulped in the process. The practice traces back to a June 2025 court ruling. Judge William Alsup found Anthropic's approach of buying physical books, cutting off the bindings, scanning the pages and training on the text was “fair use”. The $1.5 billion settlement only covered the pirated digital copies. Buying and scanning print was declared perfectly legal. That ruling handed AI labs a legal roadmap for a supply chain: buy physical books in bulk, scan them, train on the contents. ISBNdb, a database company that once helped libraries and bookshops move inventory, advertised bulk sourcing of 1,000 to one million books per order "tailored to your LLM training needs". Its pitch promised buyer anonymity because, in its own words, "the optics problem is real" and "'AI company destroys two million books' is not a headline that generates sympathy". One second-hand seller went from shifting 20 books a week to hundreds, with orders unified only by the fact that every title has an ISBN. Rare booksellers in the Netherlands, Germany and Spain report the same pattern of random titles, bulk volumes, and buyers with no interest in resale value. Nine days after the report, ISBNdb deleted the pages and denied the service ever launched. There’s an important detail worth understanding: why pre-2022 print commands a premium. It is the only text structurally guaranteed to be free of AI-generated writing, which has already seeped into the open web and newer publishing. That makes the clean pre-LLM corpus a finite asset, and the labs destroying it are also the ones actively bidding its value up. I wrote about this idea first in my newsletter: https://lnkd.in/et7jExTp And you can follow me Alex Banks for daily AI highlights and insights.
AI Copyright Law Guide
Explore top LinkedIn content from expert professionals.
-
-
When Anthropic stopped training on books that were literally pirated, they managed to hit on the one way of buying books that means no money goes to authors: buying used books. They spent tens of millions of dollars buying used books from wholesalers, in batches of tens of thousands at a time. These were shipped to Illinois, scanned, and pulped. They called this Project Panama, using a codename because they didn’t want people to know they were doing it. (It ultimately came out through court documents.) Before alighting on this plan, they were discussing licensing from book publishers, which would have meant money going to authors. But then they came up with the used books plan, and stopped all licensing discussions. Anthropic uses their huge war chest to get all the books in the world (that’s their aim) - authors get nothing. IMO there are serious questions over whether this should be legal. Yes, they are buying the books. But you can’t just do anything you like with a book once you’ve bought it. You can’t scan it and sell it as an ebook, for instance. There are limits on what you can do with books you’ve bought, where what you would be doing would compete with the book’s rights holders. As Judge Chhabria said in Meta v Kadrey, LLMs will likely compete with the books they are trained on by flooding the market. And as Dario Amodei himself said in 2021, big AI companies centralizing profits by training on books without the authors getting paid is a real concern. Whatever your view of its legality, it’s pretty clear that it sucks for authors, letting Anthropic make money at their expense. Authors should get paid when their books are used to train AI, and should have the chance to say no to that training. Anthropic’s used books strategy gives them neither.
-
AI literacy is harder than we thought. Research published this week (Ahmed et al., 2026) from Stanford & Yale shows that LLMs can extract near-verbatim copyrighted texts. Claude 3.7 Sonnet achieved 95.8% recall for some books. Gemini and Grok extracted over 70% of Harry Potter, without jailbreaking. This phenomenon is known as memorization: the encoding of specific training data in a model’s weights such that it can later be extracted in outputs. ➜ Why this matters for education. When students ask an LLM to “explain photosynthesis” or “summarize the themes in 1984,” they assume synthesis across many sources. In most cases, that's true. The problem is that they cannot tell when it is not. Research on memorization shows that, under some conditions, LLMs can reproduce long, near-verbatim passages from specific copyrighted texts. There is no signal indicating whether a response is synthesized or recalled. Under these conditions of opaque generation, attribution-based academic integrity frameworks become difficult to apply meaningfully. ➜ Consider this scenario. A student genuinely trying to learn uses AI for help. The AI produces an argument that is largely verbatim from a book the student has never seen. The student internalizes it and later writes an assignment in their own words. The work may be original in expression, but its intellectual provenance is unknowable. Who plagiarized? The student didn't know. The AI can't “know”. Intent, visibility, and traceability have collapsed. ➜ This is what students need to understand: - LLMs may reproduce specific source material rather than synthesize across sources in some cases - You cannot reliably tell whether output is memorized or transformed - Attribution becomes structurally impossible when sources are hidden AI obscures the learner’s ability to distinguish synthesis from reproduction, challenging traditional academic integrity frameworks. “Teach them to cite AI” can't be the solution. ➜ So what helps? There are no foolproof answers, but one response is to make research accountability explicit. That means teaching the slow, often invisible, work of evidence building so students can verify and defend where their ideas come from. In practice, this means introducing traceability at the level of student process: - Annotated bibliographies - Research logs documenting search strategies and decisions - Evidence tables mapping claims to sources - Explicit documentation of rejected sources - Treating process documentation as seriously as the final product - Follow-up oral presentations and peer questioning This is not a complete solution. Determined students can still game process documentation. But it shifts assessment toward intellectual work we can actually verify: source evaluation, evidence quality, the evolution of thinking across drafts, and the ability to defend choices under questioning.
-
Can Authors Keep Their Work from Being Used to Train AI Without Permission? ✍️📚🤖 If you're a writer, there's a good chance your work has already been absorbed into an AI model—without your knowledge or consent. Books, blogs, fanfiction, forums, articles… All of it has been scraped, indexed, and used to teach machines how to mimic human language. So what can authors actually do to protect their work? Here’s what’s possible (and what isn’t—yet): 🛑 Use “noAI” Clauses in Your Copyright/Terms Clearly state that your work may not be used for AI training. It won’t stop everyone, but it helps establish legal boundaries—and could matter in future lawsuits. 🔍 Avoid Platforms That Allow AI Scraping Before publishing, check the terms of service. Some platforms explicitly allow your content to be used for training; others are more protective. 🖋️ Push for Legal Reform The law hasn’t caught up to generative AI. Supporting copyright advocacy groups and legislation can help tip the scales back toward creators. 🤝 Join Opt-Out Registries Tools like haveibeentrained.com let creators see if their work was used—and request removal from certain datasets. It's not a perfect fix, but it's a start. 📣 Speak Out When authors make noise, platforms listen. Just ask the comic book artists, novelists, and journalists who’ve already triggered investigations and lawsuits. Right now, the balance of power favors the AI companies. But that doesn’t mean authors are powerless. We need visibility. Transparency. Fair compensation. And most of all—respect for the written word. Have you found your writing in an AI training dataset? What did you do? #AuthorsRights #EthicalAI #AIandWriters #GenerativeAI #Copyright #ResponsibleAI #WritingCommunity #AITrainingData #FairUseOrAbuse
-
The industry built on a legal assumption that has not yet been tested to finality. Courts are now testing it. The results will reshape how AI systems are trained, documented, and deployed. The assumption was straightforward: training a model on publicly available data is transformative. The model does not store the data. It learns patterns from it. The output is new. Fair use applies. That assumption may be directionally correct. It is not settled law. And the litigation currently in discovery will determine whether it holds, not in principle, but in the specific, detailed circumstances of what each company actually did. Courts have ruled that AI training on copyrighted books constitutes fair use, but storing pirated copies does not. That distinction is procedurally important and commercially significant. The companies that sourced training data carefully are in a different legal position from those that did not. The documentation of sourcing decisions about what was used, where it came from, and what the licensing status was at ingestion is now the evidentiary foundation of a fair use defense. Cases expanding into discovery in 2026 could compel AI companies to reveal exactly which copyrighted works were used in training, potentially costing billions and fundamentally changing how language models are developed. The structural consequence is already visible. The question of what data a model was trained on which was previously treated as a technical implementation detail is now a material legal fact. It must be documented. It must be defensible. It must survive discovery. Organizations that treated training data sourcing as an engineering decision have inherited a legal problem. The engineering happened years ago. The legal consequences are arriving now. Data provenance is not a future compliance requirement. For any organization that has deployed or built on large language models, it is a present liability question. The pipeline that built your model is now a potential exhibit.
-
Your AI reads documents like someone randomly tearing pages from a book. That's why its answers feel disconnected and strange. A new paper on Meta-chunking explains how this can be solved. When we feed documents into AI systems (for RAG), we typically chop them into bite-sized pieces. Most companies do this by arbitrary length - like cutting a book into equal pieces without caring where the scissors land. Imagine reading a strategy document where every 50 characters, someone randomly inserted a break. That's what your AI is dealing with right now. A new approach called meta-chunking changes this. Instead of random breaks, it preserves complete thoughts and ideas—just like you'd naturally read a document. Why should marketing leaders care? Because when your AI can't properly read your content, it can't: • Accurately answer customer questions • Properly understand your brand voice • Maintain consistency across responses • Connect related ideas from different documents The results are compelling: • Better performance on complex queries (with lower compute) • Works across multiple languages • Maintains a natural flow of ideas Bottom line: If you're investing in AI for customer service, content creation, or knowledge management, how your system reads and understands documents matters more than you think. Look at your current AI solutions. Are they giving fragmented, inconsistent answers? That's partially a chunking problem.
-
OpenAI lost a key discovery fight, and the message to every AI company is clear -- your training data receipts are now part of the case. A federal judge in New York ordered OpenAI to turn over internal communications about why it deleted two large datasets of pirated books known as Books1 and Books2, allegedly used to train its models. If those messages show the company knew the books were infringing and pressed ahead anyway, a good-faith fair use argument goes out the window. It also opens the door to statutory damages that can stack fast if each work is treated as a separate infringement. Courts are sending a clear signal. In multiple AI cases, they are forcing developers to open the black box, from training data to internal emails, and they are increasingly skeptical when companies try to hide behind privilege or trade secret claims. At the same time, we are seeing eye-watering settlements such as Anthropic agreeing to pay more than a billion dollars over pirated book claims. If you are building or buying AI, this is your wake up call. Governance can no longer live in a slide deck. You need to know exactly what went into your models, document consent and licenses, and assume those choices may be read aloud one day in a courtroom. What do you think this ruling means for the future of training data and fair use in AI, and how are you adjusting your own risk strategy in response?
-
In a landmark publication, the US Copyright Office (Library of Congress) has provided an exceptional case-study analysis of the use of copyrighted and intellectual property in the training of Generative AI. This publication largely affirms the protections of intellectual property in advance of existing litigation (more than 39 lawsuits) between IP rights holders (The New York Times, Thomson Reuters, The Atlantic, POLITICO ) and big tech such as OpenAI, Microsoft, Cohere, Perplexity Meta Facebook, Google The report makes a couple of key points quoted below: 1) "As noted above, some argue that the use of copyrighted works to train AI models is inherently transformative because it is not for expressive purposes. We view this argument as mistaken" 2) "Nor do we agree that AI training is inherently transformative because it is like human learning" 3) "A student could not rely on fair use to copy all the books at the library to facilitate personal education; rather, they would have to purchase or borrow a copy that was lawfully acquired, typically through a sale or license. Copyright law should not afford greater latitude for copying simply because it is done by a computer." 4) "In short, the analysis should not turn on the status of any individual entity but on the reality of whether the specific use in question serves commercial or nonprofit purposes" 5) "Copyright owners have a right to control access to their works, even if someone seeks to obtain them in order to make a fair use Gaining unlawful access therefore bears on the character of the use" 6) Factor 2 - "Where the works involved are more expressive, or previously unpublished, the second factor will disfavor fair use" 7) "Downloading works, curating them into a training dataset, and training on that dataset generally involve using all or substantially all of those works. Such wholesale taking ordinarily weighs against fair use" 8) "The speed and scale at which AI systems generate content pose a serious risk of diluting markets for works of the same kind as in their training data. That means more competition for sales of an author’s works and more difficulty for audiences in finding them" 9) "Where licensing options exist or are likely to be feasible, this consideration will disfavor fair use under the fourth factor" 10) "at this point in time, the Office recommends allowing the licensing market to continue to develop without government intervention" ForHumanity interprets the support for rightsholders that the analysis is founded upon, as positive. Highlighting that the burden to prove "fair use" is contextually driven at a minimum and an uphill battle in other contexts. We believe that supporting the robust protection of copyright is key and critical to individual and collective human flourishing as a critical motivating factor to human creativity. #aigovernance #aiethics #independentaudit #infrastructureoftrust https://lnkd.in/eTHWwGhd
-
Every day, publishers are pitched countless books on AI and how to work with it. In almost 95% of the cases, we turn them down, with good reason. AI is something that literally changes daily – evolving (some would say devolving) in leaps and bounds at a rate few can keep pace with. Now, balance that against how long it takes to get a book edited and published and you realize the issue: By the time the book finally goes to print, some of the content may well be out of date. Companies, individuals, software functions (and glitches) and entities mentioned may well have ceased to exist or fallen out of favor, and new developments will have occurred. It can take up to nine months or more for a book to go from initial idea/outline to finished product, and that’s almost an entire generation where AI is concerned. It becomes a perception problem: Even if just one aspect of the work becomes invalid, the reader will question the validity of the whole book. But if something is compelling, is it not worth rushing a book to print? Possibly, but that depends on who the author is. If you are a nationally known entity featured in the media regularly, then the book will likely sell well. If not, it’s an expensive gamble, and the odds are not in your favor. It’s also a gamble because when rushing something to print, you are skimping on quality control – something rushed to print will not have the editorial, design, and layout quality of something that goes through the measured process – and it will be crashed into the marketplace instead of being seeded carefully to retailers, distributors, foreign rights licensors, and others. AI books also have a very limited lifespan due to customer perception: Would you read a book about AI that came out even a year ago? I wouldn’t. This doesn’t mean AI is a dead topic, it just means approach it with an eye not on specifics but from “above.” AI evolves every moment, but its nature doesn’t. Its functionality changes daily, but its purpose doesn’t. Focus on the big picture that will not change, not the details that will. #publishing #books #authors #writing #Nonfiction #AI
-
The Moral Dilemma of Generative AI We are "accomplices after the fact" The world’s most powerful AI models, valued in the hundreds of billions ,were trained on data that didn’t belong to them. None of it was licensed. None of it compensated. Much of it wasn’t even disclosed. The New York Times vs. OpenAI In December 2023, The New York Times filed a lawsuit against OpenAI and Microsoft, alleging that its copyrighted journalism was used, verbatim at times, to train GPT without permission. A federal judge allowed the lawsuit to proceed on copyright grounds.¹ Authors Guild Lawsuit Well-known authors including George R.R. Martin, John Grisham, and Jodi Picoult filed suit, claiming their books were scraped and used to train large models without consent.² Meta’s Use of Shadow Libraries Court documents revealed that Meta trained its LLaMA models on datasets containing over 7 million pirated books sourced from Z-Library and Library Genesis (LibGen)—both long known as illegal repositories.³ Whistleblower Accounts Former OpenAI researcher Suchir Balaji publicly alleged that the company knowingly trained models on copyrighted data and described the practice as both unethical and economically harmful to creators.⁴ Today, 5 companies control more than 90% of the model access, compute infrastructure, and core training data pipelines.⁶ Even this post is likely hosted on infrastructure tied to OpenAI, through Microsoft’s deep integration with LinkedIn and Azure. So What’s the Point? If you use GenAI, whether directly or indirectly, you’re not just a user. We are all beneficiaries of a system built, at least in part, on stolen input. ******************************************************************************** The trick with technology is to avoid spreading darkness at the speed of light Stephen Klein is Founder & CEO of Curiouser.AI, the only AI designed to augment human intelligence. He also teaches at UC Berkeley. To learn more or sign up: curiouser.ai or contact me Hubble https://lnkd.in/gphSPv_e Sources Reuters, “Judge Allows NYT Copyright Case to Proceed” https://lnkd.in/g-zJmnd6 Wikipedia, “ChatGPT Legal Concerns” https://lnkd.in/g-MH2mwx Vanity Fair, “Meta Trained AI Using Pirated Books” https://lnkd.in/gEWW9BCi Business Insider, “OpenAI Whistleblower: Copyright Violations” https://lnkd.in/gUc7h54a arXiv, “LLM Training Data Memorization” https://lnkd.in/gqBE9SAC CSET & Epoch AI, “AI Compute and Model Concentration Report” https://lnkd.in/gyu8VPfm Cornell Law, “Definition of Accessory After the Fact” https://lnkd.in/gct65D4C