What a data quality report should contain, and how to read one sceptically

A data quality report is the document that arrives with a dataset and states what was measured, by what method, and with what result, for that delivery. It is not a datasheet, which describes why a dataset was made and how it was collected, and it is not a marketing summary. It is closer to a test report: counts with denominators, methods with versions, and a list of what failed.

This is a guide for the buyer reading one. It sets out what each section should contain, what a weak version of that section looks like, and the questions that separate a report that can be checked from one that can only be believed.

What a data quality report is for

The report has one job: to let a buyer decide whether the data is fit for the use they have in mind, using evidence they can verify against the records themselves. Documentation standards such as Datasheets for Datasets and Data Cards cover motivation, composition, collection and intended use, and they are worth asking for. The quality report is narrower and more perishable: it is about this delivery, these records, and the checks run on them.

Two properties make a report usable. Every figure has a denominator and a method, so that a duplication rate names the records it was measured over, the tier that found the duplicates and the threshold used. And every aggregate can be traced to records, so that a buyer can pick ids at random and find the per-record check results the report summarises.

Coverage, and what was dropped

Coverage says what the delivery contains against what was asked for: sources, time span, languages, record types, and for domain data the entities or repositories included. Gaps are named rather than implied: a source that was requested and could not be used, with the reason, is coverage information.

What was dropped is the mirror of coverage and the most informative section in the report. Every record that was ingested and not delivered should be counted under a reason code: duplicate of a kept record, licence unknown or not permitting the use, personal data that could not be cleared, failed extraction validation, failed verification. The ratio of dropped to delivered tells you about the source, and the distribution of reasons tells you about the pipeline. A report with no drops at all is not evidence of a clean source; it is more often evidence that nothing was checked.

Duplication rates and extraction confidence

Duplication should be reported per tier, because exact, near and semantic deduplication catch different things and cost different amounts. For each tier: the method, the threshold, the count removed, the rule for which copy was kept, and a sample of removed pairs a buyer can inspect. For a continuing delivery, duplication is also measured against earlier deliveries, since a second batch that overlaps the first is the same problem in slow motion.

Extraction confidence applies wherever fields were pulled out of unstructured text, usually by a model. The honest measures are the validation pass rate per field against the schema, the count of records that failed validation and were dropped or retried, and, separately, the agreement between the extraction and a sample that people checked by hand. A model's own confidence score is not an accuracy measurement, and a report that offers only that has not measured accuracy.

The checks each record passed

The report summarises; the records carry the detail. Every delivered record should carry the list of checks it passed, with their versions, and the report should present the same checks as counts. Typical entries, varying by domain:

The test a buyer can run in an afternoon: pick record ids at random, read their check results, and confirm that the report's counts are consistent with what the records say. A vendor that cannot support that join has a summary, not a report.

Licence summary and the evaluation split

The licence summary is the distribution of licence identifiers across the delivery, the rights basis for records that are not under an open licence, and an explicit count of records whose licence could not be determined. That last count should be zero inside the delivery, because such records belong under drops, and a summary with no unknown bucket at all has usually folded unknown into permissive.

The evaluation split section states how evaluation data was separated from training data: the unit of the split, the overlap checks run across it, the public benchmarks it was checked against, and the counts found. Splitting by source rather than by row explains why the unit matters. A report that says "held out" without a unit or a check is describing an intention.

How to read a data quality report sceptically

Five questions do most of the work:

  1. Does every figure have a denominator, a method and a version? If not, ask for them before reading further.
  2. Is anything reported as complete or perfect? Recall for personal data, licence detection and extraction accuracy are all estimates from samples, and a report that gives them without a sample size has not estimated them.
  3. Where are the drops? A delivery with none was probably not checked. A delivery with many is a candid source with a working pipeline, and the reasons tell you which.
  4. Can you trace an aggregate to records? Try it with a handful of ids before signing anything.
  5. What was not checked? The best reports say so; the others leave you to find out.

These are the same questions that run through the data questions to ask an AI vendor, applied to a single artefact, and they are cheap to ask. A report that survives them is one you can put in front of your own reviewers.

Runix Data delivers a quality report with every delivery, covering coverage, duplication rates, extraction confidence, the checks each record passed, and what was dropped and why, produced by the report stage of Runix Pipeline. Runix Data is in early access, by engagement, and the report is meant to be read this way.

Questions this raises

What is the difference between a data quality report and a datasheet?

A datasheet documents why a dataset was made, how it was collected and what it is intended for. A quality report is about one delivery: what was measured on these records, by what method, and what was dropped.

What is the most important section of a data quality report?

What was dropped and why, counted by reason. It shows whether the checks actually ran, what the source was like, and whether unknown licences or uncleared personal data were held back or quietly passed through.

How can a buyer verify a data quality report?

Pick record ids at random, read the per-record check results, and confirm the report counts are consistent with them. A report whose aggregates cannot be traced to records is a summary, not evidence.

Related to this post: Runix Data. Tell us what you are building and we reply within one business day.