Security data, with nothing live inside
Advisories, vulnerability records, logs and threat reports for models that triage, detect and explain: merged across sources, affected versions normalised, indicators defanged, and every label traceable to the rule that set it.
Part of Runix Data. This page sets out the rules we apply in this domain and the public standards they follow.
vulnerabilities CVE ID; NVD and OSV merged weaknesses CWE weakness IDs, not categories severity CVSS vector, version kept behaviour MITRE ATT&CK techniques intelligence STIX 2.1 objects indicators defanged: hxxp, example[.]com
A reference card. Each line names a public standard or convention; the rules below say how it is applied.
02 · Covers
The data this covers
Cleaned and structured from what you provide or have the rights to use, or built to a specification agreed in writing.
Vulnerability advisories
CVE records, vendor advisories and ecosystem databases for the same flaws, merged into one record each.
Logs and alerts
Endpoint, network and application events, with personal data and secrets masked.
Threat intelligence
Reports and indicators, structured and defanged so nothing in them is live.
Detection content
Rules and the events they fire on, for models that write or explain detections.
03 · Rules
The rules, and where they come from
Each rule follows a public standard or an established practice in the field, named with it, so you can check the reasoning rather than take ours on trust.
One vulnerability, one record
The same flaw appears in the NVD, vendor advisories and package ecosystems in different words. Records are merged on the CVE identifier and the aliases OSV records list, with every source kept.
Follows The CVE Program's identifiers; the NVD; the aliases field of the OSV schema.
Versions you can query
Affected and fixed versions are normalised into ranges a program can compare, per package ecosystem.
Follows The OSV schema's affected ranges; CPE names for products.
Severity as scored
CVSS vectors are kept with their version and never mixed: a v3.1 score and a v4.0 score are different measurements.
Follows FIRST's CVSS specifications.
Nothing live ships
URLs, domains and IP addresses are defanged, malware is referenced by hash rather than included, and exploit code is excluded.
Follows The defanging conventions used across threat-intelligence sharing.
Techniques named the same way
Attacker behaviour is labelled with ATT&CK technique identifiers, so two reports describing the same technique say so.
Follows MITRE ATT&CK; STIX 2.1 for structured intelligence.
Labels you can trace
Every detection label records the rule or analyst decision that assigned it, so a disputed label can be checked rather than argued about.
Follows Detection rules with stable identifiers, as in the open Sigma rule format.
04 · References
Public references
The standards and open sources these rules are built on. They are other organisations' work, linked so you can read them yourself.
- CVE ProgramIdentifiers for publicly disclosed vulnerabilities.
- National Vulnerability DatabaseNIST's analysis of CVE records, with CVSS and CPE data.
- OSVAn open vulnerability database and schema for open-source packages.
- CWEA catalogue of software and hardware weakness types.
- CVSSFIRST's standard for scoring vulnerability severity.
- MITRE ATT&CKA knowledge base of adversary tactics and techniques.
- STIX and TAXIIOASIS standards for structured threat intelligence.
- SigmaAn open, generic format for detection rules, each rule with its own identifier.
05 · Delivery
What every delivery carries
The same in every domain; the Runix Data page has the full list.
Provenance and licence, per record
Where each record came from, what was done to it, and the licence or permission it was used under.
A quality report
Coverage, duplication and the checks each record passed, plus what was dropped and why.
Evaluation kept apart
Evaluation data split from training data by source, so a score is not inflated by near-duplicates.
06 · Questions
Common questions
Can delivered data contain live malware or exploits?
No. Indicators are defanged, malware is referenced by hash, and exploit code is excluded.
How are conflicting advisories handled?
They are merged on the CVE identifier and the aliases OSV lists, with every source kept, so a model sees where the sources agree and where they do not.
Do you sell ready-made Cybersecurity datasets?
Not off the shelf. Runix Data builds to a specification agreed in writing: from data you provide or have the rights to use, or from public sources whose licences permit your use. The rules on this page apply either way.
What happens to the data we send?
It is processed only to do the work you asked for. It is not used to train models, ours or anyone else's, and it is not sold.
Send us a sample of your security data
A slice of the real data and what the model has to do with it. We reply within one business day, and the scoped plan that follows includes the parts we think are not worth doing.
Request early access