Training data provenance: what every record should carry, and why a README is not proof

Training data provenance is the record of where each training record came from, what was done to it, and under what terms it may be used. The W3C PROV definition is broader and worth keeping in mind: provenance is information about the entities, activities and people involved in producing a piece of data, used to form assessments about its quality, reliability or trustworthiness. For a training set the unit is the record, not the dataset, because the questions that matter later are asked about individual records.

This post sets out what a per-record entry should contain, why a licence detected from the file beats one copied from a README, how opt-outs are honoured, and what to do with the records whose licence cannot be determined.

Why the unit is the record, not the dataset

A dataset-level statement, such as "sourced from public repositories under permissive licences", answers none of the questions that arrive after delivery. Which records came from a source that has since withdrawn permission? Which records were transformed by a pipeline version later found to have a bug? Which records can a buyer in a given jurisdiction actually use? Each of those is a filter over records, and a filter needs a field per record to run on.

It also makes the dataset auditable rather than trusted. An auditor who can pick a record at random, follow its source identifier to the original, and reproduce the transformation is checking evidence. One who can only read a summary is reading a claim. The data questions buyers put to AI vendors are mostly requests for this evidence.

What a training data provenance entry should contain

The minimum is the set of fields that lets someone else reproduce, filter and revoke:

In practice this is a small structured block that travels with the record. An illustrative entry:

{
  "source": {"repo": "example-org/example-parser",
             "commit": "3f9c2e1", "path": "src/lexer.py"},
  "retrieved": "2026-09-14",
  "sha256": "hash-of-the-original-bytes",
  "licence": {"spdx": "MIT", "basis": "detected"},
  "transforms": ["clean", "dedup: kept from cluster c-5521", "mask"],
  "optout_checked": "2026-09-30"
}

Why a detected licence beats a README

A README describes the repository as its maintainer understood it on the day it was written. It does not describe the file that was vendored in from another project under a different licence, the file added after a relicensing, the directory that carries its own LICENSE, or the header in one file that contradicts the rest. Dual-licensed projects and notices that say "see individual files" are common enough that a repository-level label is a prior, not a finding.

Detection reads the file. Scanners such as ScanCode match licence text and headers against a database of known licences and return an SPDX identifier with a match score, per file. That has three properties a README lacks: it is per record, it is reproducible, and it disagrees out loud. When the declared and detected licences differ, the disagreement is itself a field, and the record is held until a person has looked.

The same logic applies outside code. A filings corpus carries terms of use per source, a legal corpus carries court-specific rules on reuse, and a recordings corpus carries consent per session. In each case the per-record basis is what survives an audit, which is why the domain pages, from code to legal data, each name the public standards the basis follows.

Opt-outs, and why provenance makes them executable

An opt-out is a request from a rights holder that their material not be used, made before or after a dataset was built. Public code corpora set the precedent: The Stack ships with a lookup so that a developer can check whether their code is included, and a removal process that the maintainers apply to later versions. Supplied data carries the same obligation through contract: a customer may withdraw a source, and a data subject may object.

Without per-record source identifiers, an opt-out is a promise that cannot be kept. With them it is a query: every record whose source matches, plus every derived record that lists it among its inputs, is removed, the deduplication clusters it was kept from are recomputed, and the removal is dated. The opt-out check date in the entry is what tells a buyer how current that filter was at delivery.

What to do with records whose licence is unknown

Unknown is not permissive. Copyright applies to a work whether or not a licence file is present, so a file with no detectable licence is a file with no detectable permission. The honest classes are: no licence found, a licence found but not on the SPDX list, two licences found that conflict, and a licence found that does not permit the intended use.

Each class gets the same treatment: the record is held back rather than passed through, it carries the reason as a field, and the counts per reason appear in the quality report for the delivery. Held records are not deleted; the buyer may decide to seek permission, to accept a narrower use, or to leave them out. What a buyer should not receive is a dataset in which unknown licences were silently rounded up to permissive.

Runix Data records provenance and licence for every record, builds only from data a customer provides or has the rights to use or from public sources whose licences permit the use, and reports what was dropped and why. Runix Data is in early access, by engagement, with the rules for each domain stated on its own page.

Questions this raises

Why is a licence in the README not enough?

A README describes the repository as a whole on the day it was written, while vendored files, relicensed files and per-file headers can carry different terms. Detection reads each file and returns an SPDX identifier per record, which can be reproduced and audited.

What happens to a record whose licence cannot be determined?

It is held back with the reason recorded, counted in the quality report, and left for the buyer to decide on. Unknown is treated as no permission, never rounded up to permissive.

How does provenance help with opt-outs?

A per-record source identifier turns an opt-out into a query: every record from that source and every record derived from it can be found, removed and the removal dated. Without it the opt-out is a promise that cannot be executed.

Related to this post: Runix Data. Tell us what you are building and we reply within one business day.