The AI file system for object storage
Runix FS puts a POSIX file system and a multi-tier distributed cache in front of the S3-compatible storage you already run. Training jobs, inference servers and agents read the same data through a local path, with no copies and no rewritten I/O code.
Built on Curvine, an open-source CNCF Sandbox project under the Apache 2.0 licence. Read the launch post.
# 1. Attach the bucket to the namespace $ bin/cv mount s3://training-data/imagenet \ /training-data/imagenet \ --write-type cache_mode \ -c s3.endpoint_url=https://s3.us-east-1.amazonaws.com \ -c s3.region_name=us-east-1 \ -c s3.credentials.access=$AWS_ACCESS_KEY_ID \ -c s3.credentials.secret=$AWS_SECRET_ACCESS_KEY # 2. Expose it as a local directory $ bin/curvine-fuse.sh start --mnt-path /mnt/curvine # 3. Your training code reads a path, as before $ ls /mnt/curvine/training-data/imagenet train/ val/ labels.csv
Commands from the Curvine documentation. The bucket, region and listing are placeholders.
Published results
- 5 billion
- small files in one cluster, the project's stated design capacity
- 10,000
- agent pods, each with its own volume, on 99 EKS nodes in the project's test
- 2.7×
- faster time to first token in AWS's SageMaker HyperPod reference architecture
- 9.5 GiB/s
- sequential read at 32 threads in the project's own benchmark
Figures published by the Curvine project and by AWS, not Runix service levels. Sources are linked further down this page.
Object-storage economics, local-file semantics
Object storage is where AI data lives because it is cheap and durable. It is also slow to list, slow to open and has no rename. Runix FS keeps the bucket as the source of truth and puts a file system and a cache in front of it.
POSIX, not a new API
A FUSE mount behaves like a local directory: open, read, write, rename and list. Tools such as git and inotify work against it unmodified.
A cache that tiers itself
Workers hold memory, SSD and HDD tiers. Hot data is promoted to faster tiers automatically; anything not cached is read from the bucket on demand.
Every way in
POSIX through FUSE, an S3-compatible gateway, an HDFS-compatible client for Spark and Flink, and SDKs for Java, Python and Rust, all on one namespace.
Kubernetes-native volumes
A CSI driver provisions ReadWriteMany volumes by creating a directory, not by calling a cloud API, so provisioning keeps up when thousands of pods start at once.
Your bucket stays readable
File paths map one-to-one to object keys. If the cluster is stopped, the data is still in your bucket, in its original layout, readable by anything else.
Built to stay up
Metadata is replicated across masters with Raft, cached blocks can be replicated across workers, and a built-in web UI and metrics show every component.
Where it sits in your stack
Between the workloads that read data and the bucket that keeps it. Metadata goes to the masters; data is served by the workers, which fetch from the bucket on a miss and write back to it for durability.
Workloads
Interfaces
Runix FS cluster
Your storage
Your code does not change
Four ways in, all against the same namespace. Every command below is from the Curvine documentation; hosts and bucket names are placeholders.
Warm the cache before the job starts
Load a path into the cache ahead of time, and watch the job finish.
$ bin/cv load /training-data/imagenet/train/shard-00001.tar --watch $ bin/cv fs ls /training-data/imagenet $ bin/cv report
Read it from Python
No SDK: a mounted path is a path.
from pathlib import Path
root = Path("/mnt/curvine/training-data/imagenet")
for shard in sorted(root.glob("train/*.tar"))[:2]:
print(shard.name, len(shard.read_bytes()))
Give every pod a volume
A StorageClass for the CSI driver; each claim becomes a directory.
apiVersion: storage.k8s.io/v1 kind: StorageClass metadata: name: curvine-sc provisioner: curvine volumeBindingMode: Immediate allowVolumeExpansion: true parameters: master-addrs: "m0:8995,m1:8995,m2:8995" fs-path: "/agents" path-type: "DirectoryOrCreate"
Or speak S3 to it
The gateway serves the same files to anything with an S3 client.
$ aws s3 ls s3://training-data/ \
--endpoint-url http://localhost:9900
$ aws s3 cp s3://training-data/labels.csv . \
--endpoint-url http://localhost:9900
Built for the workloads that stall on I/O
GPUs are the expensive part of an AI cluster, and they sit idle while data crosses the network one object at a time. These are the places that shows up first.
Model training
Datasets and checkpoints are cached next to the GPU nodes, so epochs after the first read from the cache instead of the bucket, and checkpoints land on a file system instead of a multipart upload.
LLM inference
A shared tier for the KV cache that outlives a single GPU's memory, so a prompt prefix computed on one replica can be reused on another.
AWS published a reference architecture for this on SageMaker HyperPod with Curvine as the shared tier, reporting up to 2.7× faster time to first token at 2,500 tokens. Read the AWS post
AI agent platforms
Every agent wants its own writable workspace. Block volumes run into per-node attachment limits long before a node runs out of CPU; a directory on a shared file system does not.
The Curvine project provisioned 10,000 volumes for 10,000 agent pods on 99 EKS nodes, all bound and running, with a storage cluster of four pods. Read the write-up
Model and artifact distribution
Weights and artifacts are pulled once from the bucket and served from the cache to every node that starts a replica, instead of each node downloading the same files.
Analytics on the lake
Spark, Flink and OLAP engines read hot tables through the HDFS-compatible client or the S3 gateway, with the cache absorbing repeated scans of the same files.
What the project has measured
These figures are published by the Curvine project from its own benchmark runs. They are not a Runix service level. During early access we measure on your data, your instances and your access pattern.
| Measure | Result | Conditions |
|---|---|---|
| Metadata operations | ||
| Create | 19,985 ops/s | 40 concurrent clients |
| Open | 60,376 ops/s | 40 concurrent clients |
| Rename | 43,009 ops/s | 40 concurrent clients |
| Delete | 39,013 ops/s | 40 concurrent clients |
| Throughput | ||
| Sequential read, 256 KiB blocks | 9.5 GiB/s | 32 threads |
| Random read, 256 KiB blocks | 7.8 GiB/s | 32 threads |
| Scale | ||
| Small files in one cluster | 5 billion | Stated design capacity |
Source: the Curvine README and benchmark documentation, which also describe the hardware and the comparison systems.
Open source, or run with us
The file system is the same. What Runix adds is the part a production team needs around it: someone to size it, deploy it, answer for it, and sign for it.
| Curvine, open source | Runix FS | |
|---|---|---|
| Licence | Apache 2.0 | Apache 2.0 core |
| Where it runs | Wherever you run it | Your AWS, Google Cloud or Azure account |
| Deployment | Self-managed | Sized and deployed with our engineers |
| Support | Community, on GitHub | Our team, 1 business day response |
| Contract | None | Runix AI Inc, MSA and DPA on request |
| Price | Free | Quoted per deployment in early access |
Common questions
How is Runix FS related to Curvine?
Runix FS is built on Curvine, an open-source distributed file system and cache released under Apache 2.0 and accepted as a Cloud Native Computing Foundation Sandbox project. Runix deploys it into your cloud account and supports it under a contract with a US company.
Where does my data live?
In your own object storage. Runix FS caches blocks on its workers, but the bucket stays the durable copy, and file paths map one-to-one to object keys, so the data stays readable without Runix FS.
Do we have to change our code?
No. Applications read and write through a FUSE mount as if it were a local directory, through the S3-compatible gateway, or through the HDFS-compatible client, so training scripts, Spark jobs and S3 tools keep working.
Which storage does it work with?
Amazon S3 and any S3-compatible store such as MinIO, plus Google Cloud Storage, Azure Blob Storage, Alibaba Cloud OSS and HDFS. The cluster runs on Kubernetes, on virtual machines or on bare metal.
How is it priced?
During early access, pricing is quoted per deployment. Tell us how much data you have and what reads it, and we come back with a quote within one business day.
Can we buy it through AWS Marketplace?
Not yet. A listing is in preparation. Until then you contract directly with Runix AI Inc and are invoiced in USD.
Put a file system in front of your bucket
Tell us where your data lives, how big it is and what reads it. We reply within one business day with a deployment plan.
Request early access