Policy & Regulation

AI and copyright: what the rulings actually turn on

Is training an AI model on copyrighted work legal? The courts have now answered enough times to see a pattern, and it isn’t the yes-or-no ruling either side wanted.

What’s emerging is a fact-specific test that turns on two things: how the data was obtained, and whether the output competes with what it was trained on.

AI copyright: training itself has survived, mostly

Two federal courts ruled in 2025 that training on copyrighted books can be fair use, which was the outcome the industry needed.

The reasoning is that training extracts statistical patterns rather than reproducing expression, so it can be transformative in the sense copyright law cares about.

That’s a significant win, and it came with conditions attached that turned out to matter more than the headline.

It also isn’t universal. A third court found training was not fair use where the output competed with the source, so the same activity has produced opposite rulings depending on the facts around it.

How you got the data is the decisive question

The largest copyright settlement on record makes the distinction unmistakable.

Anthropic agreed to pay $1.5 billion to authors and publishers, roughly $3,000 per work across about 500,000 books, with a federal judge granting final approval in July 2026, according to Norton Rose Fulbright’s litigation update.

The court found the training was fair use. It also found that downloading the books from pirate sites was not, and that second finding cost $1.5 billion.

Read that split carefully, because it’s the whole lesson. Acquisition and use are judged separately, and a lawful use of unlawfully obtained material is still an infringement.

Every decisive ruling so far has hinged on whether the data was lawfully acquired, per tracking of the 2026 cases.

Market harm is the second test

The clearest loss for an AI developer came where the product competed directly with the source it learned from.

A court found that Ross Intelligence’s use of Westlaw headnotes to train a legal AI was not fair use, focusing on market harm, because the resulting tool competed with Westlaw’s own research service.

FactorPoints toward fair usePoints against
AcquisitionLicensed, purchased, publicPirated or scraped against terms
Market effectOutput serves a different marketOutput substitutes for the source
ReproductionPatterns onlyPassages reproduced near-verbatim
What the rulings so far actually turn on.

So a general-purpose assistant trained on books sits differently from a legal research tool trained on a competitor’s legal research product. Same technique, different answer.

That distinction is doing a lot of work, and it points somewhere uncomfortable for the industry. As models get better at any given profession’s output, the market-harm argument gets stronger against them, not weaker.

A tool that summarises news competes with news. A tool that drafts contracts competes with the firms whose contracts trained it. The more capable the model, the closer it moves to the Ross facts rather than away from them.

The case everyone is watching

The New York Times sued OpenAI arguing that ChatGPT can reproduce Times articles nearly verbatim, and the case remained ongoing as of April 2026 per case tracking.

It matters because it combines both tests at once. If a model can output substantial passages from a source, that’s reproduction rather than pattern extraction, and it’s also a plausible substitute for the original.

A ruling for the Times would push far harder on the industry than the Anthropic settlement did, because settlements price a past mistake while a verdict sets a rule.

What AI copyright means if you’re building

Three practical consequences follow, and none require a legal opinion to act on.

Provenance is now a compliance artefact, not a footnote. Keep records of where training data came from and under what terms, because that’s the question courts are actually asking.

Fine-tuning on a competitor’s output is the sharpest risk in the current tests, since it hits acquisition and market harm together. And if you rely on a provider’s model, ask what indemnity they offer, because the exposure is upstream of you.

Open weights change the shape of that exposure rather than removing it. Downloading a model released under Apache 2.0 gives you a permissive licence on the weights, and says nothing about what the weights were trained on.

Data supply is tightening anyway

Legal pressure is arriving at the same moment as a supply constraint, which compounds both.

Villalobos and colleagues projected that models “will be trained on datasets roughly equal in size to the available stock of public human text data between 2026 and 2032“, and licensing is one of the few routes past that, as our piece on synthetic data covers.

Chinchilla made the arithmetic worse, finding that “for every doubling of model size the number of training tokens should also be doubled” in DeepMind’s compute-optimal work. More data is needed exactly as it becomes more expensive to obtain lawfully.

Where jurisdictions diverge

Fair use is an American doctrine, and the analysis above doesn’t transfer cleanly elsewhere.

Europe approached the question through transparency rather than through litigation, and the AI Act now requires marking generated content, as our piece on what the Act requires sets out. That’s a different lever pointed at a related problem.

Anyone operating across both is exposed to whichever regime is stricter on each question, which in practice means planning for American litigation risk and European disclosure duties simultaneously.

What would settle it

An appellate ruling on the reproduction question would do more than any settlement, because settlements deliberately avoid setting precedent.

Watch also whether licensing deals become standard practice. If frontier labs converge on paying for corpora rather than defending scraping, the legal question becomes commercially moot regardless of how the cases end.

The settlement arithmetic pushes that way already. At roughly $3,000 per work, licensing upfront looks cheap next to litigating afterwards, and finance departments tend to notice that sort of comparison.

One caution worth stating plainly. This is a live area, rulings are fact-specific, and nothing here is legal advice. Case trackers are the right place to check current status before making a decision on it.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Rundowns AI Desk

The Rundowns AI desk covers artificial intelligence research, tools, business and policy. Every factual claim we publish links to the primary source it came from, so readers can check it themselves.

Leave a Reply

Your email address will not be published. Required fields are marked *