AI & Copyright
The open questions of rights around AI training and AI output
- AI and copyright splits into three questions: is training on this material allowed, who owns the output, and how to treat an output that closely resembles an existing work.
- The three questions get weighed separately. Being fine on one doesn't mean the others are fine too.
- The answers differ by place and are still being worked out. This isn't a settled area with one answer to hand you.
- Whether rights attach to an output tends to turn on how much of a human hand went into it.
- The most solid hedge available right now is the habit of keeping a record of where material came from and how something was made.
Contents
1The analogy
Say you're putting together a single issue of a neighborhood newsletter. You want to write up last month's event, run a few photos a neighbor took, and borrow a good line you read in another newsletter.
Several separate questions come up here. Reading a lot of other people's writing and picking up a feel for it, then writing your own piece, is fine — copying a whole paragraph over directly is not, though exactly where that line sits is hard to state precisely. Running someone's photo needs their permission, and you start wondering whether even a shop sign in frame matters. On top of that, who the finished, assembled newsletter belongs to is its own open question.
The conversation around AI and copyright resembles the one happening in that little editorial room.
2In detail
The question splits three ways
The first is about material: whether it's allowed to train a model on text and images gathered from across the internet without asking permission first. The second is about output: whether a right attaches to a generated image or piece of writing, and if so, whose. The third is about resemblance: how to treat an output that ends up strikingly close to a specific existing work.
These three get weighed separately. Even where training is allowed, an output that closely resembles an original work is still a separate problem.
Carried over to the newsletter: whether it's fine to read the material is the first question, who the finished newsletter belongs to is the second, and how closely a printed line overlaps with someone else's writing is the third.
Learning from something versus copying it
One old principle keeps resurfacing here: copyright protects expression, not an idea or a style by itself. Reading a lot of work and picking up how to write from it is fine; copying someone else's sentences over directly is different.
Which of these AI training resembles is exactly what's being argued over. Viewed as a model keeping only a statistical pattern, it leans toward the first case; viewed as a process that copies and processes material at scale, it opens the second conversation. How far different places recognize a research or analysis exception varies too, so the same activity gets judged differently from place to place.
Because of that, a flat statement like "training is fine" or "training is not allowed" is usually describing one place at one moment, or presenting something unsettled as though it were already settled.
Rights on an output tend to track how much of a human hand went into it
A long-running thread in many places has been attaching rights to expression a person made. Little to no rights are often seen as attaching to an output with almost no human involvement — a short prompt, one button pressed, an image comes back.
On the other end sits work with a lot of a person's hand in it: a rough sketch, many rounds of revision, pieces chosen and refined by a person. That kind of output is more often seen as carrying rights in the parts a person actually shaped. Where "enough involvement" begins is still being worked out case by case.
The same split shows up in the newsletter: one pasted together from other people's writing and one built from original reporting get treated very differently, even though both are technically newsletters.
Resemblance gets judged on the output itself
How closely an output overlaps with a specific work is looked at separately from the training question. Even if nothing went wrong at the training stage, an image nearly identical to an existing piece still needs looking at on its own — and while the training question stays unsettled, an output different enough from any single work raises no resemblance issue.
What often gets confused here is imitating an outcome versus imitating a style. A request to generate something "in the style of" is, by the principle above, about style rather than expression, which places it outside what's typically protected — disputes tend to turn on how close the output lands to an actual work. A tool that restricts requests naming a specific style is usually enforcing its own policy, not something a law spells out.
What's still unsettled
Several questions remain open: what, if anything, should go back to creators whose material trained a model; how someone marks material off-limits for training, and who enforces that; how an output's origin should be disclosed.
The arguments run in different directions too. One side holds creators need some share coming back or fewer people will keep creating; another holds that narrowing what counts as learning makes new work harder to make. Both are arguing from the same place — wanting creative work to keep happening.
What can be done in practice right now is not treating something unsettled as though it were settled. Keeping a record of where material came from, and what tool made something and how far it went, leaves something to go back to once a standard is set. Last checked: 2026-09.
3More precisely
Keeping the terms separate keeps the conversation clear. Using material to train is a question about using material. Whether rights attach to an output is a question about recognizing a creator. Whether an output overlaps with an original work is a question about infringement. Folding all three into one word tangles the conversation, because a fact that settles one of them says nothing about the other two. Recording what an output was made from doesn't answer any of these questions by itself, but it does leave something to point back to once they're answered, which is worth more than it sounds like on a first read.
The analogy breaks down in places too. The writing and photos in a newsletter are visible and traceable back to their source; a model doesn't hold onto its training material as-is, which makes it hard to tell what actually went in just by looking at the output. A newsletter has a fixed print run, while generated output can be produced in enormous volume in very little time. And in an editorial room there's a known person to ask, while the creators whose material trained a model are typically impossible to even count, let alone locate one by one.
Last verified: 2026-09
4Try it yourself
5Common misconceptions
It's easy to think anything AI-made automatically has no copyright, but actually it tends to be judged by how much of a human hand went into it, and that standard is still being worked out.
It's easy to think that if material was used in training, the output is automatically a problem, but actually the training question and how closely an output resembles a work are judged separately.
It's easy to think imitating a style is itself an infringement, but actually style and ideas by themselves have long been treated as outside what's protected, and disputes turn on how close an output lands to a specific work.
7One-line summary
In shortAI and copyright splits into three questions — material, output, and resemblance — and because all three are still unsettled, keeping a record of what was used and how is the most solid hedge available right now.
Spotted an error or have a better analogy? Suggest an edit · Last updated2026-09-02