Datasets in the bucket, GPUs that stop waiting
For AI for Science and embodied AI groups with datasets in object storage and records that need domain judgement before a model learns from them. Runix FS mounts the bucket as a file system, cached in front of the GPUs. Runix Data applies its rules for antibody, protein and embodied AI records. Runix Models tunes and serves open-weight models.
One of six solution pages. Statuses are literal: the products named below are in early access or in development, as marked.
The stack for research labs
- Runix FS
- AI-native file system on your object storage · early access
- Runix Data
- Domain data for training and evaluation · early access
- Runix Models
- Fine-tune and serve open-weight models · early access
02 · Situation
The situation
Three things this team is usually dealing with when it calls.
The data is in a bucket, the job wants files
Sequence files, trajectories and sensor recordings sit in object storage because that is where they fit. Training code wants a path, and each framework brings its own client for the same bucket, so the GPUs wait while the data arrives.
Records where a bad merge changes the answer
Antibody and protein records come from sources that disagree on identifiers, numbering and format. Merging them carelessly changes the science, not just the formatting, and a near-duplicate on both sides of a split inflates an evaluation.
A tuned model with no baseline and no home
A model tuned on the group's data has to be compared with the model the group runs today, on a held-out set, before anyone trusts it. It then needs somewhere to run that is not a shared endpoint, with the licence on the base model checked first.
03 · Stack
The products, and the role each plays
Each one works on its own; together they are one path through the stack.
04 · Start
How it starts
A person on the other end at every step; nothing here is self-serve.
01Tell us about the data
Where it lives, roughly how big it is and what reads it, plus a sample of the records and what the model has to do with them. We reply within one business day.
02Get a scoped plan and a quote
Runix FS is sized and deployed with you in your own AWS, Google Cloud or Azure account and quoted per deployment. Data is quoted per project or by volume and Models per engagement, both before any work starts.
03Mount, train, evaluate
The bucket keeps its layout and stays independently readable, so nothing is migrated. Training and evaluation sets are split by source, and the tuned model is measured against the one you run today before it is served.
05 · In writing
What is in writing
The parts a review asks for, stated once and linked to the page that holds them.
Data use
The data you provide to Runix Models trains and evaluates the model you commission, and nothing else. Inputs to Runix Data are data you provide or have the rights to use, or public sources whose licences permit your use.
Where it runs
Runix FS runs in your AWS, Google Cloud or Azure account, supported under contract with Runix AI Inc. Model serving is single-tenant, in your cloud account or on capacity arranged per engagement; an MSA and a DPA are available on request.
Statuses, literally
Runix FS is in early access, per deployment; Runix Data and Runix Models are in early access, by engagement. Performance figures on the FS page are the Curvine project's own measurements, labelled as such, not a Runix service level.
06 · Reading
Reading for this team
Engineering notes that cover the mechanics behind this page.
- Introducing Runix FS: an AI file system for object storage
- From Raw to Model-Ready: A Practical Guide to AI Data Pipelines
- Rebuild AI Unix: why Runix Lab is building an operating system for AI
Other teams:
07 · Questions
Common questions
Do we have to move the data out of the bucket?
No. Paths map one-to-one to object keys, and the bucket keeps its layout and stays independently readable. In cache mode the bucket is the source of truth and writes pass through; in file-system mode operations complete in the cluster and then sync to the bucket.
Which scientific data do you handle?
Within AI for Science, biological data: antibody and protein records. Embodied AI is a separate domain with its own rules. For anything else, describe the data and we will say plainly whether we have the judgement for it.
Can the tuned model run on our own cluster?
Serving is dedicated and single-tenant, in your cloud account or on capacity arranged per engagement, behind an OpenAI-compatible API. Weights and access are set in the engagement contract before training starts.
Tell us where the data lives
Tell us what you run today and what you are trying to change. We reply within one business day with a concrete plan, and we say which parts we would not do.
Talk to sales