Fair Use Protected the Training. It Doesn't Protect Your Output.
Every favorable AI training ruling gets read in marketing meetings as permission to publish. It isn't. The defense belongs to the party that did the training, and the risk in a published output belongs to whoever published it. Understanding where the two questions separate is most of the practical work.
The Conflation That Creates the Exposure
Fair use is a defense to copyright infringement. Defenses are personal to a defendant and specific to an act. When a court considers whether ingesting a corpus of books to train a model is fair use, the act under examination is the model developer's copying. The defendant is the model developer. The holding does not transfer to you, and it does not describe your conduct.
Your act is different: you generated something and put it on a website, in an ad, in a shipped product. The question for that act is the ordinary one — does the thing you published copy protectable expression from a copyrighted work? That analysis starts at substantial similarity, and it does not care how the model was trained.
- Defendant: the model developer
- Turns on transformativeness and market harm
- Separately, on how the copies were obtained
- Resolved in litigation you are not a party to
- Defendant: you, the publisher
- Turns on substantial similarity to a specific work
- Unaffected by the training ruling
- Resolved by a demand letter addressed to you
What the Four Factors Are Actually Testing
The statutory factors are familiar; how they behave in AI disputes is less so. In practice, two of the four are close to dispositive.
The pattern that has emerged across the first wave of decisions is not "AI training is fair use" or "AI training is infringement." It is narrower and more useful: the more the system substitutes for the original in its own market, and the less lawful the acquisition of the copies, the worse it goes. A tool that summarizes what a document says is on very different ground from one that outputs a competing version of the document.
Acquisition Is Its Own Exposure
One of the most commercially relevant distinctions to come out of this litigation is between using a work to train and obtaining the copy in the first place. A holding that the training use was transformative does not retroactively legitimize a library assembled from pirated files; the downloading can be its own infringement with its own damages. For a buyer, this converts an abstract legal debate into a concrete diligence item: where did the training corpus come from, and will the vendor say so in writing?
Pre-Publication Checks That Actually Reduce Risk
None of this argues against using generative tools. It argues for a small number of checks placed where the risk concentrates — on outputs that are public, commercial, and close to someone's protected expression.
The Asymmetry Worth Planning Around
Machine-generated expression is not protectable as your copyright, but it can still infringe someone else's. You carry the downside without the upside. That asymmetry is a reason to be deliberate about which assets you generate: high-volume, low-stakes content where copying by competitors is irrelevant is a good fit; brand marks, flagship creative, and anything you would want to enforce against a copycat is not.
Frequently Asked Questions
A court held that training on copyrighted books was fair use. Doesn't that settle it?
It settles one question for one defendant on one record. It does not address whether an output you publish is substantially similar to a protected work, and it does not bind courts in other circuits or on different facts. Treat favorable training rulings as reducing vendor risk, not as clearing your publication risk.
Which fair use factors matter most in AI cases?
Factor one, on whether the use is transformative, and factor four, on market effect. Factor three behaves unusually because training typically uses whole works, which courts have accepted where the copying serves the transformative purpose. Factor two rarely moves the outcome in mass-ingestion cases.
Does it matter that our use is commercial?
It is relevant but not decisive. Commerciality weighs against fair use under factor one, yet plenty of commercial uses are fair. What tends to matter more is whether your use substitutes for the original in its own market — a commercial use that serves a different purpose fares better than a non-commercial one that displaces the source.
How was the training data acquired, and why should we care?
Because unlawful acquisition is a separate wrong from the training use, and a favorable ruling on the latter does not cure the former. Ask vendors to describe corpus provenance and to represent that copies were lawfully obtained. Refusal to put it in writing is itself informative.
Can we rely on our vendor's copyright indemnity?
Only within its exclusions and its cap. Common carve-outs cover disabled filters, user-supplied inputs, prompts naming third-party content or styles, modified outputs, and outdated model versions. Map your actual workflow against the exclusion list; teams often discover their standard process falls outside coverage.
Can we copyright the content our AI produces?
Only the human-authored contribution — selection, arrangement, and substantive modification. Purely machine-generated elements are not registrable. Assets you intend to enforce against copying should have a documented human authorship story, or should be commissioned rather than generated.
Put the Review Where the Risk Is
The training-data question will be resolved over years, in courts, by parties that are not you. The output question is resolved every time someone hits publish. A similarity check on public-facing assets and a rule against naming creators in prompts cost almost nothing and remove most of the realistic exposure.
Spend the legal budget on vendor terms and provenance representations. Spend the process budget on the last step before publication.