Executive Summary
Ownership of AI training data sounds settled until two contracts disagree. Legal teams answer one way and model providers answer another, and the gap decides who carries the risk. As AI budgets climb, that gap turns into real liability.
Ownership is a bundle of rights, not a single fact. Copyright, database rights, trade secret protection, and a contractual license each behave differently, and the license is usually the one that matters. Provenance decides the rest. Hugging Face hosts hundreds of thousands of datasets whose cards range from precise to meaningless, and a Creative Commons tag means nothing if the uploader never held the right. Indemnity will not close the gap. Most coverage defends outputs rather than inputs, and the caps land at a few million dollars. Watch the lineage requirement. Demand the dataset list, a data processing addendum that blocks reuse of your inputs, and weights you could host yourself.
Ask your legal team who owns the data you feed into a model and you get a confident answer. Ask the model provider the same question and the answer shifts. That gap is the story. Ownership sounds like a settled fact. In practice it is a negotiation.
Most enterprises assume ownership travels with the data. They paid to collect it, cleaned it, labeled it, stored it. So they assume it is theirs outright. Then a provider contract grants a broad license to use that data, and the assumption dies quietly in an appendix nobody read.
Ownership is not one thing. It is several rights bundled together. Copyright in the content is one. The database right is another. Trade secret protection is a third. A contractual license is the fourth, and often the only one that matters in practice. When a vendor says it owns your data, ask which of these it means.
Provenance makes it worse. Training corpora mix scraped text, licensed content, synthetic output, and customer uploads. Each source carries different rights. Hugging Face hosts hundreds of thousands of datasets with labels that range from precise to meaningless. A dataset card that says CC-BY may cover the packaging and not the underlying words. Creative Commons licenses attach to a work, and only the rights holder can attach them. If the uploader never held the right to license the text, the license is noise.
None of this is new law. It is newly expensive. A company that licensed a photo archive for marketing never imagined those images would land in a training set. Publishers, broadcasters, and stock libraries have all filed suit, and the outcomes remain open. The buyers caught in the middle are the ones who cannot show where their data came from.
Owning the Data Is Not Owning the Model
Here is the distinction buyers miss. You can own every byte of your corpus and own none of the model trained on it. Fine-tuning does not transfer ownership of the base weights. It produces a new artifact governed by the base model’s license. Reading that license takes ten minutes. Anthropic and its peers each write different terms for what you may do with derivatives.
So a claim that you own your model is usually false. What you own is a license to use weights or an endpoint. That license may forbid hosting the model for third parties, training a competing model, or moving the weights to another cloud. Ownership of data buys you less than the invoice suggests.
The open-weight world complicates this further. A model you download under a public license is still governed by that license. Some terms cap commercial use. Some require attribution. Some grant the provider rights over your improvements. Read them before you build a product on top. A permissive-sounding name does not make the terms permissive.
Indemnity Covers You Until It Does Not
Providers now offer IP indemnification, which promises to defend you if a model’s output infringes someone’s copyright. Microsoft Azure and other hyperscalers bundle it at enterprise tiers. Read the scope before you relax. Most indemnities cover outputs, not inputs. If you fine-tuned on data you had no right to use, you carry that risk alone.
The caps matter too. A promise capped at a few million dollars is small against a class action. Indemnity also tends to exclude the exact scenarios that scare buyers, such as training data scraped from the open web.
Indemnity is also narrow. The provider agrees to defend you against a copyright claim on generated text. It does not agree to replace your lost revenue, your customer trust, or the engineering time you burn answering lawyers. Those costs never appear in the clause.
Demand the Lineage Before You Sign
Get the provenance of any model you depend on. Ask which datasets trained it and under what licenses. Ask whether your inputs train the provider’s next model, and get the answer in writing. Ask what happens to your fine-tuned weights if you leave. Then ask for the export format.
Push for a data processing addendum that names your training inputs and forbids their reuse. Push for weights you could host yourself if the relationship ends. If a provider refuses both, treat that refusal as information. A vendor that will not document its lineage will not defend you when the lineage is challenged.
C2PA content credentials and newer provenance standards point toward a future where lineage is machine readable. Model Context Protocol does something similar for tool and data access. Neither is a contract yet. Until they are, the paper is the product.
Ownership is not a feeling. It is an enumerated right. If the contract does not name the right you need, you do not have it. Read the appendix, because that is where your data actually lives.
Related reading. who owns the data your AI model learns from, and why the cost of using AI is not falling.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data centre economics. No vendor spin.
