Antibody and protein data, where a bad merge changes the answer
Sequence and structure records for models in antibody and protein work: formats validated, one numbering scheme applied across sources, accessions reconciled, and measured values kept apart from computed ones. Our scope here is biological data.
Part of Runix Data. This page sets out the rules we apply in this domain and the public standards they follow.
sequences FASTA; alphabet and length checked numbering one scheme: IMGT, Kabat, Chothia structures PDB mmCIF, chains mapped identity UniProt accessions reconciled antibodies SAbDab and OAS conventions repertoires AIRR Community standards
A reference card. Each line names a public standard or convention; the rules below say how it is applied.
02 · Covers
The data this covers
Cleaned and structured from what you provide or have the rights to use, or built to a specification agreed in writing.
Antibody sequences
Heavy and light chains, paired where the source pairs them, numbered under one scheme.
Protein sequences
Sequences with their accession, organism and the evidence behind them.
Structures
Experimental structures with their chains mapped to the sequences they contain.
Assay measurements
Binding and activity values with their units and method, never mixed with predictions.
03 · Rules
The rules, and where they come from
Each rule follows a public standard or an established practice in the field, named with it, so you can check the reasoning rather than take ours on trust.
Formats validated, not just parsed
Sequences are checked for their alphabet, length and truncation, not only read. A file that parses is not necessarily a sequence that makes sense.
Follows The FASTA format and the validation built into public sequence databases.
One numbering scheme
IMGT, Kabat and Chothia number antibody residues differently and draw the loops in different places. One scheme is applied across every source, so a position means the same thing everywhere.
Follows IMGT numbering and tools such as ANARCI that number antibody sequences.
Accessions reconciled
The same protein appears under different identifiers in sequence databases, structure archives and papers. They are mapped to one entity, and conflicts are reported rather than settled silently.
Follows UniProt accessions; cross-references between UniProt and the PDB.
Measured, not asserted
Measured values keep their units and method, and computed properties are labelled as computed. A binding site that was predicted is not recorded as one that was observed.
Follows Structural antibody databases such as SAbDab, which record each structure's experimental method and resolution and, where available, a curated binding affinity.
Scope stated
Our work in this domain is limited to biological data: antibody and protein records. Materials, climate and chemistry data are outside it.
Follows A stated scope, so what is outside it is clear.
04 · References
Public references
The standards and open sources these rules are built on. They are other organisations' work, linked so you can read them yourself.
- UniProtThe reference database of protein sequences and function.
- RCSB Protein Data BankThe US data centre of the PDB archive of experimentally determined structures; computed models on the site are labelled as such.
- IMGTThe international ImMunoGeneTics information system and its numbering.
- ANARCIAntibody numbering and receptor classification.
- SAbDabThe Structural Antibody Database (now SAbDab2), from Oxford's OPIG.
- Observed Antibody SpaceA database of antibody repertoire sequences.
- AIRR CommunityStandards for adaptive immune receptor repertoire data.
05 · Delivery
What every delivery carries
The same in every domain; the Runix Data page has the full list.
Provenance and licence, per record
Where each record came from, what was done to it, and the licence or permission it was used under.
A quality report
Coverage, duplication and the checks each record passed, plus what was dropped and why.
Evaluation kept apart
Evaluation data split from training data by source, so a score is not inflated by near-duplicates.
06 · Questions
Common questions
Do you work on materials or chemistry data?
No. Our scope here is biological data: antibody and protein records.
Where can we see how you approach this data?
We teach it in a free course, released under CC BY-SA 4.0 and taught in Chinese, with an English overview on ai4s.runixcloud.io.
Do you sell ready-made AI for Science datasets?
Not off the shelf. Runix Data builds to a specification agreed in writing: from data you provide or have the rights to use, or from public sources whose licences permit your use. The rules on this page apply either way.
What happens to the data we send?
It is processed only to do the work you asked for. It is not used to train models, ours or anyone else's, and it is not sold.
Send us a sample of your antibody or protein data
A slice of the real data and what the model has to do with it. We reply within one business day, and the scoped plan that follows includes the parts we think are not worth doing.
Request early access