Listen to the article
In 1999, a college dropout named Shawn Fanning and his friend Sean Parker released a piece of software called Napster that let strangers swap music files over the internet. Within two years, it had roughly 80 million registered users — one of the fastest adoption curves for any software in history at its time. Within three, it was dead.
It didn’t die alone. There was Aimster. Grokster. Morpheus. Kazaa. An entire business model for mass-market file-sharing was gradually and then quickly wiped out by a string of court cases in the 2000s culminating in a blow at the Supreme Court in MGM Studios, Inc. v. Grokster. The instrument that dealt the blow was copyright law, wielded by record companies fighting for what they saw as their survival in the digital age.
Now, a bigger, more diverse, better-resourced group of copyright-holders has their eyes set on another threat: the generative AI industry.
A tidal wave
The scale of lawsuits being brought against AI companies — and the names attached to them — has been staggering. In September 2023, the Authors Guild and a collection of authors including John Grisham, George R.R. Martin, and Jodi Picoult filed a class-action lawsuit accusing OpenAI of copyright infringement, seeking statutory damages of up to $150,000 for each infringed work. The authors alleged that OpenAI “copied Plaintiffs’ works wholesale, without permission or consideration,” then fed those copies into large language models, which used the copyrighted text for “pretraining” in which a massive collection of text teaches the model the relationships between words and other artifacts of language. The complaint called that process, which is the essential basis for building an AI model, “systematic theft on a mass scale.” Like Napster before it, the plaintiffs understood the entire enterprise to be premised on copyright theft.
Other industries followed suit. The next month music giants Universal Music Group, Concord, and ABKCO filed a complaint against Anthropic alleging that the training of AI models on artists’ lyrics constituted theft, and they’ve since filed an additional suit for over $3 billion in potential statutory damages. That December, The New York Times hit OpenAI and Microsoft with a multi-billion dollar lawsuit seeking damages for scraping millions of copyrighted news articles to build a product they said would both “substitute for The Times and steal audiences away from it.” And in 2025, Disney and Universal Studios brought Hollywood into the fray, suing Midjourney over its image generator, seeking up to $150,000 for every willfully infringed work. The studios called Midjourney’s models a “bottomless pit of plagiarism.”
What has FIRE been doing in the AI space?
How did FIRE become a leading voice on AI? By defending free speech in legislatures, courts, research, and emerging technology.
Read More
The wave of lawsuits — almost all of which are still pending — is not simply a story of industries trying to strangle a rival technology or recoup lost revenues. In many ways, traditional creative industries see themselves where the record industry stood in 1999: staring at a technology that could make their business obsolete. Like the internet and all that came with it, AI has the potential to transform how we send and receive content and information. What traditional creative industries’ role will be after that transformation, should it happen, is uncertain. Will people still check news sites, or just ask an AI who has digested the news? Will we read books, or ask an AI for a bespoke book in the style of an author we like? While those questions remain open and debated, it’s only natural that rights-holders find themselves in a defensive posture, looking to ensure their copyrighted works don’t potentially feed the “engines of their own destruction,” as the Authors Guild’s complaint put it.
Both sides believe the stakes are existential for their industries, and each argues the future of expression depends on its theory of copyright prevailing. These lawsuits strike at the heart of what copyright law is for and its role in the creative ecosystem.
So: What is copyright law for?
Copyright is supposed to serve free expression
“Congress shall make no law . . . abridging the freedom of speech.” If there’s a lesson from this series so far, it is the extent to which that line transcends boundaries of form and technology. Any and all restrictions on expression must have a strong justification. Copyright law is no exception.
It derives its authority by living alongside the First Amendment in the Constitution. Article I grants Congress the power to secure exclusive rights to authors “to promote the Progress of Science and useful Arts.” This premise has not only been seen by courts as aligned with the First Amendment, but intended by the Framers to work at its service. As the Supreme Court wrote in 1985’s Harper & Row v. Nation Enterprises, copyright was meant to be “the engine of free expression” by providing the “economic incentive to create and disseminate ideas.”
Still, there’s a clear tension. The enforcement of copyright inherently requires government policing of expression through the court system to be effective, and much of that expression could be valuable from the standpoint of First Amendment values. Courts and Congress have responded by drawing important limiting principles meant to keep copyright in service of free expression. These “built-in First Amendment accommodations” codified in the Copyright Act of 1976 include the idea-expression distinction, first established in the 1879 Supreme Court case Baker v. Selden. Under that distinction, copyright is limited only to particular instances of expression — a specific movie like Star Wars, for example, or a specific character like Mickey Mouse. By contrast, copyright does not extend to general ideas, facts, principles, and themes — think a scientific theory like general relativity or a fiction trope like enemies-to-lovers. This distinction is intended to keep copyright law from limiting the free flow of information and knowledge.
If the courts embrace a theory that has the opposite effect — one that diminishes our ability to create and share new speech — we’d risk losing in copyright law what has been a powerful tool for encouraging expression to the side of censorship.
Another First Amendment accommodation in the Act is the fair use doctrine, which allows unlicensed use of copyrighted material under certain conditions through balancing the rights of authors against the need to advance free expression. Fair use lies at the core of copyright law’s evolution through new technologies and the fight over AI training. We’ll give it a deeper look momentarily.
But first, we can summarize generally that these limiting principles act as important safeguards to protect freedom of speech. When these safeguards fail, copyright can become a tool of speech suppression — and it has a track record as one. Litigants have wielded infringement claims to suppress leaked documents, unflattering footage, and critical commentary. With the digital age, bogus content takedown notices under the Digital Millennium Copyright Act routinely knock criticism, parody, and journalism offline. This gives free speech advocates a key stake in defending First Amendment safeguards like fair use in copyright law, to make sure the system works as intended and serves free expression rather than censorship.
What is fair use?
Justice Joseph Story provided a rough sketch of the principles of fair use in 1841’s Folsom v. Marsh, making the critical distinction between using copyrighted passages from a book for “fair and reasonable criticism,” and use that would substitute or “supersede the use of the original work.” The defendant had copied extensive portions of a multivolume collection of George Washington’s writings into a shorter competing biography. Story concluded that the new work displaced demand for the original, establishing an early distinction between transformative use and market substitution.
Story’s opinion laid out a brief protocol for distinguishing legitimate use from infringement, one Congress would later codify in Section 107 of the Copyright Act as a four-factor test. Under that now-statutory test, courts weighing a fair use claim are to consider:
(1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
(2) the nature of the copyrighted work;
(3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
(4) the effect of the use upon the potential market for or value of the copyrighted work.
Examples that have traditionally supported a finding of fair use include quoting brief portions of a book in a review that comments on it, using clips from a movie or television series for a reaction video to provide criticism or commentary, making personal copies of a work for private consumption, and parody.
One of the best-known parody cases involved 2 Live Crew’s version of Roy Orbison’s “Oh, Pretty Woman” — the center of 1994’s Campbell v. Acuff-Rose Music, Inc. In that case, the Supreme Court held that the fact 2 Live Crew sold its parody commercially didn’t prevent it from being a fair use.
More importantly, Campbell reshaped the first factor of the fair use inquiry by emphasizing whether the challenged use was “transformative.” The Court asked whether the defendant used the copyrighted work for a new purpose or character — adding new expression, meaning, or message — or instead merely “supersede[d] the objects” of the original by functioning as a substitute for it, like Upham’s George Washington biography. That framework would become especially important as courts confronted technologies that required copying to enable uses copyright law had never previously considered.
Fair use and new technology
An early and consequential example of fair use confronting emerging technologies was the Supreme Court’s consideration of the VCR in 1984’s Sony v. Universal. Hollywood argued the VCR was an infringement machine, allowing people to keep personal copies of the entirety of any content aired on television. The Supreme Court disagreed, noting users were simply using the technology to view content they had been invited to freely view, but on their own schedules rather than at the designated air time. Counter to the movie industry’s expectations, studios went on to make fortunes from the home video market driven by the technology they had tried to strangle in the crib.
Other courts have confronted whole-work copying that was instrumental in nature — using information gleaned from copyrighted works to create transformative new products. In Sega v. Accolade and Sony Computer Entertainment v. Connectix, the Ninth Circuit held that developers could copy entire software programs for reverse engineering when that copying was necessary to understand unprotected elements (in this case, the rules, procedures, and information needed to build software and platforms compatible with Sony and Sega products) and carried out for the purposes of a transformative use. Together with Sony v. Universal, these decisions establish that the third fair-use factor does not prevent complete copying — as we saw, the transformative purpose can require analysis of the whole work.
Courts applied the same principle to online indices and search engines. In Authors Guild v. Google, the Second Circuit held that Google could scan millions of books in full to create a searchable index, a use the court found “highly transformative” even though it involved wholesale copying of entire works without permission. Similarly, courts have permitted image-search engines to copy and display reduced-resolution thumbnails because the copies served a new function — helping users locate information — rather than simply replacing the original photographs.
The application of copyright to AI should result in more speech rather than less.
The Supreme Court echoed the logic of the circuit courts in Google v. Oracle, holding that Google’s limited copying of Java’s Application Programming Interface (API) for building the Android platform was fair use. Google had copied parts of the API — which “allows programmers to call upon prewritten computing tasks for use in their own programs” — so programmers could use commands familiar to them when building smartphone apps for Android. The Court reasoned that Android was a transformative new platform and not a substitute for Java SE, that the copied code was largely functional (i.e., a method for operating the system rather than a form of creative expression), and that Google copied only what was needed to let programmers carry their existing knowledge of Java’s commands into the new Android environment. Above all, the Court emphasized copyright law should not become a barrier to the creative progress that is “the basic constitutional objective of copyright itself.”
The common theme in these cases is courts asking whether the “copying” of copyrighted material was used to generate new expression, new knowledge, or new public value — and ruling in defendants’ favor when they find that it has. This contrasts with the copying seen in the Napster line of cases and to some extent in Justice Story’s original consideration of George Washington biographies, where the copyrighted work was merely redistributed with little if any transformative elements.
Applying the fair use test to AI training
So where does AI training fall?
In June 2025, we received some strong signals. Two federal judges in the Northern District of California — Judge William Alsup in Bartz v. Anthropic and Judge Vince Chhabria in Kadrey v. Meta — issued the first major rulings applying the fair use test to generative AI training. Both cases involved a collection of authors accusing AI companies of copyright infringement for creating copies of their work and training AI models on them.
On the first three factors, the two judges reached a shared conclusion.
Applying the first factor — and the test established in Campbell v. Acuff-Rose Music, Inc. — both judges found the purpose and character of using copyrighted works to train a general-purpose AI model to be transformative. “Spectacularly so,” in the words of Judge Alsup.
It starts with how AI training works. Here’s Judge Alsup’s description:
Anthropic used copies of Authors’ copyrighted works to iteratively map statistical relationships between every text-fragment and every sequence of text-fragments so that a completed LLM could receive new text inputs and return new text outputs as if it were a human reading prompts and writing responses.
Let’s put that in simpler terms. The training process exposes a model to enormous quantities of text so that it can learn patterns in how words and concepts relate. Those patterns are reflected across billions of numerical parameters known as model weights. Those weights are not necessarily numerical representations of the content it is trained on, but rather a representation of what it learned about language from reading the text. For example, repeated exposure to sentences associating “speech” with positive values will cause training to adjust many model weights throughout the matrix defining how the model understands the relationship between “speech” and other words. As a result, when prompted about speech, the model may become more likely to describe it in positive terms. The reality is a bit more complex, but this gives you a rough idea what the millions of numerical parameters which inform the response to any given prompt look like.
Much like the information gleaned from reverse engineering Sega’s copyrighted software, the patterns an AI model learns are not necessarily copyrighted works. Copyright protects an author’s particular expression, not the broader relationships among words, concepts, themes, or writing techniques that readers — or even AI models — may learn and deploy from it.
We need stronger First Amendment guardrails on copyright, not weaker ones.
But even when Judge Alsup accepted, for purposes of his ruling, the authors’ claim that the models had in some sense “memorized” copyright-protected elements of expression like stories and characters, he found that memorization didn’t change the “purpose or character” of using those works for training (i.e., to generate new expression).
Judge Alsup’s reasoning drew on the plaintiffs’ own analogy between AI training and teaching a person to read and write, an analogy which was built to show the copies were used as intended — to be read and learned from — rather than for a transformative new purpose. This analogy bore implications which cut strongly against the plaintiffs. As Judge Alsup put it, “if someone were to read all the modern-day classics because of their exceptional expression, memorize them, and then emulate a blend of their best writing, would that violate the Copyright Act? Of course not.” Copyright law has never required a reader to receive a license each time they recall a book or apply something they learned from it.
Unpersuaded by the plaintiffs’ arguments about AI models’ inputs, both judges focused on the outputs: whether infringing content made its way to readers. On the records before them, the judges found the evidence for that lacking. Judge Alsup wrote that Claude created “no exact copy, nor any substantial knock-off. Nothing traceable to Authors’ works.” Judge Chhabria likewise found that Meta’s AI model Llama “cannot currently be used to read or otherwise meaningfully access the plaintiffs’ books,” even when researchers used prompts designed to make the model reproduce copyrighted works.
The models also used output controls intended to reduce the risk that users could obtain infringing passages. Judge Alsup likened this to Google Books imposing limits on how many snippets of text a user could see from any book in its preview feature, “preventing its search tool from devolving into a reading tool.” Had Claude delivered infringing passages to users, he emphasized, the authors “would have a different case.”
So what was the copyrighted information used for, if not to reproduce copyrighted content? Judge Alsup stated it plainly: “Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different.” Likewise, seeing the novelty in what AI companies were doing with the copyrighted materials wasn’t a particularly tough call for Judge Chhabria: “There is no serious question that Meta’s use of the plaintiffs’ books had a ‘further purpose’ and ‘different character’ than the books.” While plaintiffs’ books were used for entertainment or education, he continued, Meta’s AI models could be used to edit emails, translate languages, write code, or any number of tasks: a brand new expressive product. This was “highly transformative.”
The second and third factors largely followed the path laid by earlier technology cases.
The second factor favored the authors because the works selected for training were “highly expressive” in nature — placing it “closer to the core of intended copyright protection” per Campbell than, say, the functional computer code in the Java API case — and were chosen for their expressive qualities. But this was not very significant by itself. Judge Alsup described the second factor’s principal role as helping courts assess the purpose and amount of the copying, while Judge Chhabria noted that it has “rarely played a significant role” in the ultimate fair-use determination.
The third factor, by contrast, favored the AI companies. Like the VCR, Google Books, and search engine cases, the copying of entire works was accepted if it was reasonably necessary to accomplish the transformative purpose. In the case of AI, as Judge Chhabria put it, “feeding a whole book to an LLM does more to train it than would feeding it only half of that book.”
The fourth factor: Market impact
The judges diverged most sharply on the fourth factor, which considers “the effect of the use upon the potential market for or value of the copyrighted work.” Even more than the first factor, the Supreme Court at one point called it “the single most important element of fair use.” And in many ways, this factor captures the central stakes of the fight over AI training: To what extent is AI a legitimate competitor to traditional creative industries?
The factor also puts free-speech advocates in a tough position. On one hand, it’s important to preserve incentives to create. But on the other, it’s difficult for a free speech advocate to accept barriers to entry into the marketplace of expression. In the face of this conundrum, the judges adopted different approaches.
They agreed on direct substitution, the traditional fourth factor concern: whether the defendant’s use gives the public a replacement for the copyrighted work. As with the first factor, both judges viewed the theory as weak because neither system meaningfully reproduced the plaintiffs’ books such that they acted as substitutes. But the judges diverged on a broader and more novel question: Can AI harm authors in a copyright-relevant sense without reproducing their works at all?
The plaintiffs argued that models trained on copyrighted books could generate an enormous volume of competing expression, depressing the market for human-created works — including “by creating alternative summaries of factual events, alternative examples of compelling writing about fictional events” — even when no individual output was substantially similar to any one book. Judge Alsup considered this lawful competition from new expression, analogizing the claim to a complaint that “training schoolchildren to write well would result in an explosion of competing works.” The Copyright Act, in his view, “seeks to advance original works of authorship, not to protect authors against competition.”
Judge Chhabria took the concern more seriously. Generative AI, he reasoned, is unlike the creation of a single secondary work (i.e., a work utilizing copyrighted material) because it can produce “literally millions of secondary works, with a minuscule fraction of the time and creativity used to create the original works it was trained on.” In his view, that scale makes the big picture market substitution analysis highly relevant. AI-generated works could flood the same markets as those for specific human-created works and crowd out lesser-known authors even if well-known ones survive. He pointed to particularly vulnerable markets: An LLM capable of producing accurate information about current events, for example, could “greatly harm” the print news market. Similarly, books on subjects like cooking and gardening could be displaced by an LLM’s advice.
Judge Chhabria disagreed with his colleague as to whether this represented legitimate competition outside the purview of copyright law. He emphasized that the models compete better precisely because they were trained on the creative expression in the books they now compete against — feeding the “engines of their own destruction,” as the Authors Guild complaint put it.
And this indirect form of substitution was still substitution. Judge Chhabria explained that a reader who buys an AI-generated romance novel instead of a human-written one has replaced the latter with the former, and a world where that happens at scale is a world with lost sales and less incentive to write original works — cutting against the core purpose of copyright law.
But the plaintiffs never connected Llama’s capabilities to likely harm in the markets for their particular books, leading Judge Chhabria to rule for Meta in the absence of that evidence. Still, he warned that “it seems likely that market dilution will often cause plaintiffs to decisively win the fourth factor — and thus win the fair use question overall — in cases like this.”
Before moving on, there are a couple points of clarification worth making. First, the market impact of AI training wasn’t the only cause for a split. Judge Alsup ruled Anthropic’s use of pirated copies — as opposed to the lawfully obtained copies considered in the analysis outlined above — was not fair use. This led to Anthropic paying out a $1.5 billion settlement to the plaintiffs to compensate for the 482,000 books which they had pirated for their use in AI training. Meta had more luck. Judge Chhabria’s emphasis on the market impact factor gave him little reason to treat lawful copies differently than pirated copies: “The loss of isolated sales to AI developers is not the kind of market harm that could tip the scales for the plaintiffs.”
Second, these are just the opinions of two district court judges. All the lawsuits brought by plaintiffs we detailed in the introduction — from John Grisham to Disney to The New York Times — remain pending, and it could be a while before we reach any definitive guidance from the higher courts. Still, the split between Judges Alsup and Chhabria on the fourth factor is a helpful indicator of one place the legal debate might center going forward.
Where does the First Amendment point?
So which way do First Amendment values point in this particular dispute? It can be a tough call. Underlying that call are questions where principled free speech advocates could reach different conclusions, and they speak to some of the same questions that have always made copyright a divisive subject: Can limiting the expression of one group be justified by incentivizing the expression of another group? From a constitutional standpoint, is human-created work more important to incentivize than AI-created work? Is the training process in any sense “theft?” Some questions are beyond the pay grade of a free speech advocate, and others, like what effects new technology will have on the market for expression, are anyone’s guess.
History has given us ample reason for caution in predicting the future. New technologies often look most threatening before their markets mature. The VCR was cast as an existential threat to Hollywood before becoming a major source of revenue. The internet devastated parts of the record industry’s old business model in part because of software like Napster, yet streaming later helped drive the industry to a period of incredible growth. So predictions of market harm may prove either prescient or badly mistaken.
In the face of this uncertainty, we’re left emphasizing traditional First Amendment concerns, and there are a number.
Chief among them is the classic fear of power over expression concentrating in a few hands. The plaintiffs’ theory of market dilution, if taken up and enforced by courts, would hand incumbent industries a veto over a process critical to a whole medium of expressive technology. And however the market for expression would have ultimately shaken out without interference, the immediate effect of accepting indirect dilution could be to let established industries wipe out or hamstring a new medium of expression. And unlike the Napster cases, this wouldn’t be for the act of redistributing copyrighted works free of charge, but because it learned from their works.
This risks weakening a key First Amendment safeguard on copyright law: the idea-expression distinction. Recall that the distinction is that copyright covers an author’s particular expression, not the underlying facts, ideas, and concepts. The plaintiffs’ theory puts AI-assisted creation in the infringing category because it learns relationships between words from their works, something plaintiffs themselves analogized to human reading and learning. Accepting this as a basis for activating the market dilution analysis would extend property rights over particular expression into the building blocks of language — ideas about grammar, structure, and sequencing — embedded in the expression, potentially hindering the use of those building blocks to create transformative new expression. And again, learning from existing ideas to produce new ones is what both the marketplace of ideas and the fair use doctrine seek to facilitate and protect.
How does the First Amendment apply to AI regulation in hiring and health care?
Governments are regulating AI used in hiring and health care. But when does regulating decision-making tools become regulating speech?
Read More
Weakening First Amendment safeguards as AI becomes a powerful communications medium is learning the exact wrong lessons from the internet — where systems like notice-and-takedown implemented to protect copyright in the wake of a new technology too often worked by removing speech first and working out the details later (frequently to the detriment of fair use). We worry similarly about copyright being turned into a potent tool for AI censorship where public figures, corporations, and other powerful actors can weaponize the deference of automated copyright-abiding systems to police disparaging outputs. We need stronger First Amendment guardrails on copyright, not weaker ones.
Those guardrails sustain the ultimate principle giving copyright its place under the First Amendment — its design to promote the creation and spread of expression. It is a means, not an end. The theory contemplated by Judge Chhabria risks suppressing a new expressive medium for the aim of insulating existing markets from a forecast of existential threat that may prove mistaken. In fact, like the VCR, the incumbent creative industries may ultimately find AI lowers the cost of creation and bolsters their industries rather than rendering them obsolete.
So does training AI on large datasets fall under fair use and First Amendment protection? We’ve received some early evidence that it belongs in the same fair use tradition as the VCR, the search engine, and digital libraries, even while there might be a dispute on the market impact factor. Courts will have to resolve those questions. Whatever answers emerge, the First Amendment interest will be the same: The application of copyright to AI should result in more speech rather than less. That means more knowledge created, more ideas exchanged, more methods of expression in more hands. If the courts embrace a theory that has the opposite effect — one that diminishes our ability to create and share new speech — we’d risk losing in copyright law what has been a powerful tool for encouraging expression to the side of censorship.
Read the full article here
Fact Checker
Verify the accuracy of this article using AI-powered analysis and real-time sources.

