AI Policy, Law & Safety · AI Copyright & Intellectual Property
What Is 'Fair Use' and How Does It Apply to AI Training Data?
Fair use is a US legal doctrine allowing limited use of copyrighted material without permission under certain circumstances, weighed through factors like purpose, nature of the work, amount used, and market effect; AI companies commonly invoke it to justify training on copyrighted content, but whether that argument holds up is still being actively contested and decided case by case in court.
Legal disclaimer
This page provides general information only and is not legal advice. Laws vary by jurisdiction and change over time. Consult a licensed attorney in your jurisdiction before making decisions based on this content.
Key takeaways
- Fair use is a defense within US copyright law that permits certain unauthorized uses of copyrighted material, evaluated through a multi-factor balancing test rather than a fixed rule.
- The four traditional factors generally considered are the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the market for the original.
- AI companies frequently argue that training a model is 'transformative' — serving a different purpose than the original work — which is a key consideration under the first factor.
- Rights holders counter that large-scale copying for commercial AI training, and potential market substitution by AI-generated outputs, weighs against a fair use finding.
- Fair use is a distinctly American legal doctrine; other countries have their own, sometimes different, copyright exceptions that may apply differently to AI training.
A Legal Balancing Test, Not a Simple Rule
Fair use is a doctrine within US copyright law that allows certain uses of copyrighted material without needing permission from the rights holder, under specific circumstances. It’s not a blanket exception, and it doesn’t work like a fixed checklist where meeting one condition automatically qualifies a use. Instead, fair use operates as a balancing test: courts weigh several factors together, and the outcome depends heavily on the specific facts of each situation. This flexibility is intentional — fair use was designed to adapt to new kinds of uses that lawmakers couldn’t have anticipated in advance, which is exactly why it’s become central to the debate over AI training data.
Because fair use is decided case by case rather than through a fixed formula, it produces genuine uncertainty. Two AI companies using seemingly similar training practices could, in principle, see different legal outcomes depending on details like what data they used, how they obtained it, and how their models are ultimately deployed.
The Factors Courts Weigh, and How They Map to AI
Fair use analysis typically considers four factors: the purpose and character of the use, including whether it’s commercial and whether it’s “transformative”; the nature of the copyrighted work itself; the amount and substantiality of the portion used; and the effect of the use on the potential market for the original work.
AI companies training on copyrighted material generally lean hardest on the first factor, arguing that training a model is transformative — the goal isn’t to republish or compete with the original book or article as a substitute for reading it, but to extract statistical patterns that help build a general-purpose system serving an entirely different function. This transformative-use argument has succeeded in other technology contexts in the past, such as certain search engine and indexing cases, which is part of why AI companies find it an appealing framework.
Rights holders challenge this on multiple fronts. They point out that AI training typically involves copying entire works, not small excerpts, which cuts against the “amount used” factor. They also argue that AI-generated outputs can end up substituting for the original works in the marketplace — an AI system that can write in an author’s style, or summarize a paywalled article’s content, could reduce demand for the original in ways that weigh against fair use under the “market effect” factor, often considered one of the most influential factors in the overall analysis.
Why This Debate Doesn’t Have One Universal Answer
Because fair use is fact-specific, the same general question — “is AI training fair use?” — can have different answers depending on which AI company, which training dataset, which specific copyrighted works, and which resulting AI outputs are being examined. A model trained on a narrow, licensed dataset used strictly for internal research might present a very different fair use analysis than a model trained on large amounts of scraped copyrighted journalism and later offered as a commercial product that can reproduce similar content. This is exactly why multiple separate lawsuits against different AI companies are working through courts rather than a single case settling the matter for the entire industry at once.
Bottom Line
Fair use is a flexible, multi-factor legal doctrine — not a simple yes/no rule — and while AI companies commonly argue that training on copyrighted material is transformative and therefore fair use, rights holders push back on grounds like wholesale copying and market harm. Courts are actively working through these arguments case by case, so there is no single settled answer covering all AI training practices.
Go deeper
Important caveats
- Fair use determinations are made case by case by courts, so no single blanket rule currently confirms or denies that AI training is always fair use.
- This is general information, not legal advice; fair use analysis is fact-specific and outcomes can vary significantly.
Frequently asked questions
Is fair use a right or a defense?
Fair use functions as a legal defense against a claim of copyright infringement, rather than a right you can invoke in advance. Someone using copyrighted material can be sued, and then argue fair use as a defense, with a court ultimately deciding whether the use qualifies.
Does fair use exist outside the United States?
The specific doctrine of 'fair use' as structured under US law is largely a US concept. Other countries have their own copyright exceptions, sometimes called 'fair dealing' or similar terms, which can differ meaningfully in scope from the US fair use test.
Does using copyrighted material for a noncommercial purpose automatically make it fair use?
No. Noncommercial purpose is one consideration that can weigh in favor of fair use, but it isn't determinative on its own. Courts weigh multiple factors together, and commercial AI training has still been argued as potentially transformative despite its commercial nature.
Related questions
- Is It Legal to Train AI Models on Copyrighted Books and Articles?
- Can You Copyright Something an AI Helped You Write?
- Who Owns the Output of an AI Image Generator?
- Can ai companies be compelled to disclose their training data sources?
- Can an AI Be Listed as an Inventor on a Patent?
- Do AI Image Generators Train on Copyrighted Art?
Sources
- [1]US Copyright Office — United States Copyright Office
Written by Editorial Team
Last updated July 25, 2026
Get one well-sourced answer a week
No spam. Unsubscribe anytime.