Data you can audit, a model you can measure
For teams that train or tune models and have to show where the data came from and whether the result beats what they run today. Runix Data builds training and evaluation sets with provenance per record. Runix Models fine-tunes open-weight models and serves them single-tenant. Runix FS reads it all from the bucket you already have.
One of six solution pages. Statuses are literal: the products named below are in early access or in development, as marked.
The stack for model teams
- Runix Data
- Domain data for training and evaluation · early access
- Runix Pipeline
- The tooling that makes data model-ready · in development
- Runix Models
- Fine-tune and serve open-weight models · early access
- Runix FS
- AI-native file system on your object storage · early access
02 · Situation
The situation
Three things this team is usually dealing with when it calls.
Raw data that nobody can vouch for
Source material arrives duplicated, half-structured and with no record of its licence. A near-duplicate that lands on both sides of a split inflates the score, and nobody notices until the model ships. Cleaning it is a project of its own, and the rules differ by field.
A fine-tune with no baseline
A tuned model is only worth deploying if it beats the model you run today on your task. Without a held-out set and a before-and-after comparison, the decision is a demo and a feeling. Licence terms on the base model are checked last, if at all.
Training jobs that wait on storage
The datasets live in object storage because that is where they are cheap to keep. Each framework then brings its own way of reading the bucket, and the GPUs idle while the data arrives.
03 · Stack
The products, and the role each plays
Each one works on its own; together they are one path through the stack.
04 · Start
How it starts
A person on the other end at every step; nothing here is self-serve.
01Send a sample and the task
A slice of the real data, not a description of it, and what the model has to do with the result: train on it, be evaluated on it, or both. We reply within one business day.
02Get a scoped plan and a quote
Data engagements are priced per project or by volume and quoted before any work starts; the plan says which checks each record has to pass and what we think is not worth doing. Models engagements are quoted after scoping and before any training starts, with weights and access set in the contract.
03Train, evaluate, then serve
Your data trains and evaluates the model you commission, and nothing else. Serving is dedicated and single-tenant, in your cloud account or on capacity arranged per engagement; Runix FS, if you use it, is sized and deployed with you.
05 · In writing
What is in writing
The parts a review asks for, stated once and linked to the page that holds them.
Data use
The data you provide trains and evaluates the model you commission, and nothing else. Inputs are data you provide or have the rights to use, or public sources whose licences permit your use.
Weights, access and pricing
Weights and access are set in the engagement contract before training starts. Data is priced per project or by volume and Models per engagement, both quoted first; an MSA and a DPA are available on request.
Statuses, literally
Runix Data and Runix Models are in early access, by engagement, and Runix FS is in early access, per deployment. Runix Pipeline is in development with design partners.
06 · Reading
Reading for this team
Engineering notes that cover the mechanics behind this page.
- From Raw to Model-Ready: A Practical Guide to AI Data Pipelines
- How to tell whether a model change made things worse
- Introducing Runix FS: an AI file system for object storage
Other teams:
07 · Questions
Common questions
Do you clean our data or build new data?
Both. Runix Data cleans and structures data you provide or have the rights to use, and builds task data, such as coding tasks verified by running their tests, to a scope agreed before work starts.
How do you know the tuned model is better?
It is evaluated on a task-specific held-out set, before and after, against the model you run today. Quantised variants are served only where that evaluation shows quality holds.
Which base models can you tune?
Open-weight families whose licences permit your use, for example Qwen, Llama, Mistral, Gemma and DeepSeek. The licence is checked before any training starts.
Send a sample, and the task the model has to do
Tell us what you run today and what you are trying to change. We reply within one business day with a concrete plan, and we say which parts we would not do.
Talk to sales