Layer 02 · Data: see the whole stackRunix Data · Early access

Code data, verified by running it

Repository and task data for code models and coding agents, the focus of Runix Data: licence-checked file by file, deduplicated across forks, scrubbed of secrets, and verified by executing the tests rather than reading them.

Part of Runix Data. This page sets out the rules we apply in this domain and the public standards they follow.

code: what each record is checked against
licence     SPDX identifier, detected per file
provenance  repository, commit, path
secrets     key and token patterns, entropy
duplicates  exact hash + MinHash near-dup
tasks       fail-to-pass, pass-to-pass
evaluation  split by repository, decontaminated

A reference card. Each line names a public standard or convention; the rules below say how it is applied.

02 · Covers

The data this covers

Cleaned and structured from what you provide or have the rights to use, or built to a specification agreed in writing.

Source code corpora

Files with their repository, commit and licence, for pre-training and continued training of code models.

Repository tasks

An issue, the patch that resolves it and the tests that prove it, for training and evaluating coding agents.

Review pairs

A diff and the review comments it drew, for models that review code rather than write it.

Evaluation sets

Tasks held back from training and checked against public benchmarks, so a score measures ability rather than memory.

03 · Rules

The rules, and where they come from

Each rule follows a public standard or an established practice in the field, named with it, so you can check the reasoning rather than take ours on trust.

Licence first

Every file carries a licence identifier detected by scanning the file, not trusted from a README. Files whose licence does not permit your use are left out, and the licence travels with every record that stays.

Follows The SPDX License List for identifiers; scanners such as ScanCode; public code corpora such as The Stack, built from permissively licensed repositories with an opt-out.

One copy of each file

Forks, vendored dependencies and generated or minified files make up much of raw code. Exact hashing removes copies, and near-duplicate detection catches the files that differ by a header or a whitespace change.

Follows The MinHash near-deduplication used to build large public code corpora.

No secrets leave

API keys, tokens, private keys and passwords are found with pattern and entropy rules and removed or replaced before anything leaves the pipeline.

Follows The rule sets of open-source scanners such as gitleaks and detect-secrets.

Run, not read

A task counts only if, in a clean container, the reference patch makes its target tests pass, the unpatched repository makes those tests fail, and the rest of the suite still passes. A task whose target tests pass without the patch is dropped.

Follows SWE-bench's execution-based evaluation, with fail-to-pass and pass-to-pass tests.

Statements that match the tests

A task statement describes what the tests actually check, so a model is not graded on behaviour the statement never asked for.

Follows The human review that produced SWE-bench Verified, which screened tasks for underspecified statements and unfair tests.

Evaluation kept clean

Evaluation tasks are split from training data by repository, not by row, and checked for overlap with public benchmarks.

Follows Standard decontamination practice for code benchmarks.

04 · References

Public references

The standards and open sources these rules are built on. They are other organisations' work, linked so you can read them yourself.

  • SPDX License ListStandard short identifiers for licences used in open-source and other shared software and data.
  • ScanCode ToolkitOpen-source licence and copyright detection.
  • The Stack (BigCode)A public code corpus of permissively licensed source, with an opt-out.
  • SWE-benchExecution-based evaluation of coding agents on real repository issues.
  • gitleaksOpen-source secret scanning for repositories.
  • detect-secretsOpen-source detection of secrets in code.

05 · Delivery

What every delivery carries

The same in every domain; the Runix Data page has the full list.

Provenance and licence, per record

Where each record came from, what was done to it, and the licence or permission it was used under.

A quality report

Coverage, duplication and the checks each record passed, plus what was dropped and why.

Evaluation kept apart

Evaluation data split from training data by source, so a score is not inflated by near-duplicates.

06 · Questions

Common questions

Which licences do you accept?

Permissive licences by default. Anything else only when you ask for it and the licence allows your use, and the licence identifier travels with every file either way.

How do you know a coding task is fair?

It is run: in a clean container, the reference patch must make its target tests pass, the unpatched repository must make those tests fail, and the rest of the suite must still pass. The statement is then checked against what the tests actually test.

Do you sell ready-made Code datasets?

Not off the shelf. Runix Data builds to a specification agreed in writing: from data you provide or have the rights to use, or from public sources whose licences permit your use. The rules on this page apply either way.

What happens to the data we send?

It is processed only to do the work you asked for. It is not used to train models, ours or anyone else's, and it is not sold.

Send us a sample of your code tasks

A slice of the real data and what the model has to do with it. We reply within one business day, and the scoped plan that follows includes the parts we think are not worth doing.

Request early access