Runix Pipeline is managed data preparation: we turn raw, duplicated, half-structured source material into something a model can read. It is delivered as an engagement, not self-serve software — this page explains what an engagement looks like so you know what to expect before you write to us. Pipeline is being built with design partners, so early projects get direct engineering attention.
1. Scoping
Every engagement starts with your data and your goal. You describe the sources (documents, exports, crawls, databases), roughly how much there is, and what the output feeds — analytics, retrieval, or model training. We come back with a written scope: stages to run, delivery format, timeline and price. Scope is agreed before work starts and changes are agreed in writing.
2. The five stages
- Ingestion — pulling from your sources into a working store, with provenance kept per record.
- Cleaning — format normalisation, error and noise reduction.
- Deduplication — exact, near-duplicate and semantic tiers, applied to the depth the use case needs.
- Structured extraction — schema-validated extraction from unstructured text, so downstream code gets fields, not prose.
- PII masking — personal information identified and masked before data travels; the masking step fails closed.
Not every project needs every stage — the scope names which ones run and why. The background reasoning is in Building AI data pipelines.
3. Delivery
You receive the processed dataset in the agreed format — dataset files, database tables or an API handoff — together with a quality report: what was ingested, what was dropped and why, duplication rates, extraction validation results and masking coverage. The report is part of the deliverable, not an extra.
4. Starting a project
Tell us about your data — sources, volume, and what it needs to feed. We reply within one business day; if the fit is right, the next step is a written scope.