Training data provenance is the record of where each training record came from, what was done to it, and under what terms it may be used. The W3C PROV definition is broader and worth keeping in mind: provenance is information about the entities, activities and people involved in producing a piece of data, used to form assessments about its quality, reliability or trustworthiness. For a training set the unit is the record, not the dataset, because the questions that matter later are asked about individual records.
This post sets out what a per-record entry should contain, why a licence detected from the file beats one copied from a README, how opt-outs are honoured, and what to do with the records whose licence cannot be determined.
Why the unit is the record, not the dataset
A dataset-level statement, such as "sourced from public repositories under permissive licences", answers none of the questions that arrive after delivery. Which records came from a source that has since withdrawn permission? Which records were transformed by a pipeline version later found to have a bug? Which records can a buyer in a given jurisdiction actually use? Each of those is a filter over records, and a filter needs a field per record to run on.
It also makes the dataset auditable rather than trusted. An auditor who can pick a record at random, follow its source identifier to the original, and reproduce the transformation is checking evidence. One who can only read a summary is reading a claim. The data questions buyers put to AI vendors are mostly requests for this evidence.
What a training data provenance entry should contain
The minimum is the set of fields that lets someone else reproduce, filter and revoke:
- Source identifier. Precise enough to fetch the original again: a repository, commit and path for code; a URL and its canonical form for web text; a document id and page range for filings; a session id for recordings.
- Retrieval date and method. When the source was read and how, because sources change and permissions change with them.
- A hash of the original. So that a later copy can be shown to be the same bytes, or not.
- Transformations applied. Each stage that touched the record, with the tool and version: cleaning, deduplication and the cluster it was kept from, extraction, masking.
- Licence identifier. An identifier from the SPDX License List, such as MIT or Apache-2.0, with how it was determined: detected from the file, declared by the repository, or asserted by the party that supplied the data.
- Rights basis. The licence, a contract, or the supplier's own rights: the reason this record may be used for this purpose at all.
- Opt-out check. When the source was last checked against the removal requests that apply to it.
In practice this is a small structured block that travels with the record. An illustrative entry:
{
"source": {"repo": "example-org/example-parser",
"commit": "3f9c2e1", "path": "src/lexer.py"},
"retrieved": "2026-09-14",
"sha256": "hash-of-the-original-bytes",
"licence": {"spdx": "MIT", "basis": "detected"},
"transforms": ["clean", "dedup: kept from cluster c-5521", "mask"],
"optout_checked": "2026-09-30"
}
Why a detected licence beats a README
A README describes the repository as its maintainer understood it on the day it was written. It does not describe the file that was vendored in from another project under a different licence, the file added after a relicensing, the directory that carries its own LICENSE, or the header in one file that contradicts the rest. Dual-licensed projects and notices that say "see individual files" are common enough that a repository-level label is a prior, not a finding.
Detection reads the file. Scanners such as ScanCode match licence text and headers against a database of known licences and return an SPDX identifier with a match score, per file. That has three properties a README lacks: it is per record, it is reproducible, and it disagrees out loud. When the declared and detected licences differ, the disagreement is itself a field, and the record is held until a person has looked.
The same logic applies outside code. A filings corpus carries terms of use per source, a legal corpus carries court-specific rules on reuse, and a recordings corpus carries consent per session. In each case the per-record basis is what survives an audit, which is why the domain pages, from code to legal data, each name the public standards the basis follows.
Opt-outs, and why provenance makes them executable
An opt-out is a request from a rights holder that their material not be used, made before or after a dataset was built. Public code corpora set the precedent: The Stack ships with a lookup so that a developer can check whether their code is included, and a removal process that the maintainers apply to later versions. Supplied data carries the same obligation through contract: a customer may withdraw a source, and a data subject may object.
Without per-record source identifiers, an opt-out is a promise that cannot be kept. With them it is a query: every record whose source matches, plus every derived record that lists it among its inputs, is removed, the deduplication clusters it was kept from are recomputed, and the removal is dated. The opt-out check date in the entry is what tells a buyer how current that filter was at delivery.
What to do with records whose licence is unknown
Unknown is not permissive. Copyright applies to a work whether or not a licence file is present, so a file with no detectable licence is a file with no detectable permission. The honest classes are: no licence found, a licence found but not on the SPDX list, two licences found that conflict, and a licence found that does not permit the intended use.
Each class gets the same treatment: the record is held back rather than passed through, it carries the reason as a field, and the counts per reason appear in the quality report for the delivery. Held records are not deleted; the buyer may decide to seek permission, to accept a narrower use, or to leave them out. What a buyer should not receive is a dataset in which unknown licences were silently rounded up to permissive.
Runix Data records provenance and licence for every record, builds only from data a customer provides or has the rights to use or from public sources whose licences permit the use, and reports what was dropped and why. Runix Data is in early access, by engagement, with the rules for each domain stated on its own page.
Questions this raises
Why is a licence in the README not enough?
A README describes the repository as a whole on the day it was written, while vendored files, relicensed files and per-file headers can carry different terms. Detection reads each file and returns an SPDX identifier per record, which can be reproduced and audited.
What happens to a record whose licence cannot be determined?
It is held back with the reason recorded, counted in the quality report, and left for the buyer to decide on. Unknown is treated as no permission, never rounded up to permissive.
How does provenance help with opt-outs?
A per-record source identifier turns an opt-out into a query: every record from that source and every record derived from it can be found, removed and the removal dated. Without it the opt-out is a promise that cannot be executed.