Every file system in front of a bucket has to answer one question before any other: when the mount and the bucket disagree, which one is right? There are two honest answers, and they lead to two different behaviours. In cache mode the bucket is the only truth and the cluster holds copies; in file-system mode the cluster is the truth and the bucket follows behind. This post describes both as the Curvine project documents them, which workloads fit which, and what to ask a vendor whose page does not say.
What source of truth means for a mount
A source of truth is the place whose contents win: if a file exists there, it exists; if it was deleted there, it is gone; if two copies differ, that one is correct. For a plain object store the answer is trivial, because there is only the bucket. Once workers hold blocks, a metadata service holds a namespace and a client holds a write buffer, a file can be in several places, and the system has to declare which one counts.
That declaration decides everything a user will notice: how long a write takes to acknowledge, whether another tool reading the bucket sees the new file, what a delete through the mount does, and what is lost if the cluster disappears. None of it shows in a benchmark, which is why it is the part of a product page most worth reading twice.
AI file system cache mode: the bucket is the only truth
In cache mode the cluster never holds the only copy of anything. A write through the mount passes through to the bucket and is acknowledged on the bucket's terms; the cached copy is invalidated, and the next read fetches the new bytes from the bucket. A read that misses goes to the bucket, fills the cache and returns, and a delete through the mount deletes the object in the bucket, because there is nowhere else for the file to be.
The consequences follow directly. Anything else that reads the bucket, another cluster, a pipeline, a person with a console, sees the same state the mount sees, give or take the bucket's own consistency. The reverse is not automatic: a cached copy can be stale after another tool changes the object in the bucket, so a mount that must see external writes is configured to check the object's size and modification time against the bucket before each read. Losing the whole cluster loses nothing but cached copies, and there is no sync to monitor, because nothing is ever behind.
What you pay is the bucket's write path on every write: a small write costs an object put, and a rename costs what a rename costs on an S3-style store, a copy and a delete per object. Cache mode is a read accelerator with a pass-through write path, and should be judged as one.
File-system mode: the cluster is the truth, the bucket follows
In file-system mode an operation completes inside the cluster first and is acknowledged there. Writes land on the workers' cache tiers and the metadata service records them; a background process then syncs the result to the bucket. The bucket is eventually consistent with the cluster, where eventually means the sync lag under current load. In Curvine this is the default.
This is what makes the file system behave like one. A rename is a metadata change in the cluster; on an S3-style store it becomes a delete plus a full copy, and the Curvine documentation warns against renaming large directories on such backends because the copy takes a long time, blocks subsequent operations and costs storage and bandwidth. Small and partial writes are absorbed by the cache and reach the bucket as whole objects when the sync runs; a checkpoint is written at cache speed and synced afterwards.
The costs are the mirror image. Between the acknowledgement and the sync the bucket is behind, and any tool reading it directly sees old data or no file. Unsynced data depends on the cluster for its durability, so how the cluster protects it, by replicating metadata with Raft and by whatever the configuration does for data blocks, is a durability question rather than a performance one. And changes made to the bucket by other tools are not seen by the mount until an explicit resync, because the cluster, not the bucket, is the truth.
Which workloads fit which mode
The split is mostly about who writes and who else reads.
- Cache mode fits data that is produced elsewhere and read many times: training corpora curated by a pipeline, evaluation sets, model weights served to inference nodes. It also fits data where the bucket is the system of record and no copy may exist that the bucket does not know about.
- File-system mode fits workloads that write a lot and reread what they wrote: checkpoints, agent workspaces with one directory per agent, intermediate outputs of a preprocessing job, logs. It fits when nothing else writes to the bucket, whose role is then backup and hand-off rather than live shared state.
- Both at once is the normal shape of a training cluster: the dataset mounted in cache mode, the output directory in file-system mode. The useful property is that the choice is per mount, not per cluster.
The wrong choice fails quietly: a dataset in file-system mode that a pipeline keeps updating in the bucket trains on stale samples until someone resyncs, and checkpoints in cache mode write at bucket speed and turn every save into a pause. Neither produces an error.
Questions to ask any AI file system about where the truth lives
A vendor's page may describe POSIX, caching and tiers without stating the consistency model. These questions get to it, and belong next to the data-handling questions a procurement review already asks.
- Which is the default mode, and can it differ per mount or per path?
- When a write is acknowledged, where is the data: in the bucket, in the cluster, or in a client buffer?
- What does a delete through the mount do to the bucket, and when?
- If the cluster is lost, what unsynced data goes with it, and how is it protected until the sync?
- How does the mount learn about objects written to the bucket by something else?
- What does a rename cost in the bucket, and when is that cost paid?
- What happens when a sync fails: retries, an alert, or silence?
- Does the bucket keep its layout, so that it stays readable without the cluster?
Answers of the form "it depends on the mode" are the right ones, provided the vendor then says which mode; a product that answers with a throughput figure has answered a different question. The glossary defines the terms, and the introduction to Runix FS gives our own answers.
Runix FS, built on Curvine, offers both modes: cache mode, where the bucket is the source of truth and writes pass through, and file-system mode, where operations complete in the cluster and then sync to the bucket in the background, with paths mapping one-to-one to object keys in either case. It is in early access, deployed per customer in their own cloud account, and we publish no performance figures of our own for either mode; the quick start brings up a single-node cluster, attaches a bucket and mounts it.
Questions this raises
What is cache mode in an AI file system?
A mode in which the bucket is the only source of truth: writes pass through to the bucket before they are acknowledged, reads that miss the cache fetch from the bucket, and a delete through the mount deletes the object in the bucket.
What is the difference between cache mode and fs mode?
In cache mode the bucket is authoritative and the cluster only holds copies. In fs mode operations complete in the cluster first and sync to the bucket asynchronously, so the bucket is eventually consistent and external changes to it need an explicit resync.
Which mode should training data use?
Data produced elsewhere and read many times fits cache mode, so that other tools and the pipeline see one consistent bucket. Outputs such as checkpoints and agent workspaces fit file-system mode, where writes complete in the cluster and sync later.