Object storage and file systems get used interchangeably in architecture diagrams, and the confusion becomes expensive once a training job or an agent platform reads from a bucket through code written for files. This post defines both, lists what a bucket cannot do, explains what a POSIX file system on object storage has to add, and is honest about what it still cannot promise.
What object storage actually is
An object store is a key-value service for large values. You put an object under a key, get it back by that key, list the keys that share a prefix, and delete it. The key is a string; the slashes in it are a naming convention, not a directory tree.
Each object is written whole: a put replaces the entire value under that key, and there is no call that changes bytes in the middle of an existing object. Multipart upload assembles a large object from parts, but the result is still one immutable value until the next put replaces it.
In exchange you get durability, capacity that grows without planning, and a price per gigabyte a file server cannot match, which is why AI data lands in buckets and why a file system layer in front of the bucket exists as a category at all.
What a file system promises that a bucket does not
The POSIX interface, specified by the Open Group and IEEE, is what almost every data loader and shell tool was written against. Its guarantees are small and specific, and each one is missing from an object API.
- Rename is a metadata operation. Moving a file or a directory changes a name, not the data. On a general purpose S3 bucket a rename is a copy followed by a delete, and renaming a prefix is one copy and one delete per object under it.
- Partial writes exist. A process can seek to an offset and write a few bytes, append to a log, or truncate a file; a general purpose S3 bucket has none of these, and the nearest equivalent is to read the object, change it in memory and put it back.
- Directories are real. A directory can be empty, is listed cheaply, and carries its own permissions and timestamps; a prefix is just the set of keys that start with the same string.
- Listing is local. A readdir on a local file system reads a few blocks; listing a prefix is a paged network call, and a prefix with millions of keys takes many pages, each a round trip.
- Consistency is defined across operations. A write followed by a read on the same host sees the new bytes, and two processes coordinate with locks. Amazon S3 documents strong read-after-write consistency for single objects (Amazon S3 user guide), but there is still no lock, no atomic multi-object operation and no notion of an open file.
What a POSIX file system on object storage has to add
A layer that presents a bucket as a file system cannot change what the bucket does; it has to add three components of its own and keep them in agreement.
A metadata service. Something has to hold the directory tree, sizes, timestamps, permissions and the mapping from a path to the blocks or objects that hold its bytes. Once that service exists, rename and directory listing become operations on it rather than on the bucket, which is where the speed comes from. It also has to survive a node failure, which is why it is usually replicated with a consensus protocol such as Raft.
A cache. Every open, read and stat that reaches the bucket pays a network round trip. A data cache on local memory and disk turns repeated reads into local reads, and a write buffer turns small writes into whole objects, which is what makes partial writes possible at all.
Consistency rules. This is the part vendors describe least and buyers should ask about most. The layer has to say where the truth lives: either the bucket is the source of truth and every write passes through to it before it is acknowledged, or the cluster is the source of truth and the bucket is updated afterwards. The second is faster, and it means the bucket is briefly, or indefinitely, behind, which decides what another tool reading the bucket directly sees.
What it still cannot do
A file system layer removes the cost of the missing operations; it does not remove their nature. Keep these limits in mind when a product page says POSIX without qualification.
- A rename that is instant in the metadata service is still a copy and a delete when it reaches an S3-style bucket; if the mount syncs later, a large directory rename takes a long time to settle there.
- A write that completes in the cluster first is not durable in the bucket until the sync runs, and a tool reading the bucket directly may see stale data or no file at all.
- Changes made to the bucket by other tools are invisible to the metadata service until something tells it to look again. Paths that map one-to-one to object keys make that resync possible; they do not make it automatic.
- Byte-range locks, atomic append from many writers and hard links are either emulated or absent. Workloads that depend on them, such as a database file, belong on block storage.
- The object store's own limits remain: request rates per prefix, latency on a miss, and the egress bill when the cache is cold.
Questions to ask about any POSIX layer
The claim of POSIX on S3 covers everything from a thin FUSE adapter that turns each read into a GET to a distributed file system with its own metadata and cache. The way to tell them apart is to ask five questions.
- Where does metadata live, and what happens to it if that node dies?
- What does rename do, in the cluster and in the bucket?
- When is a write durable, and in which place?
- What does a delete through the mount do to the bucket?
- If another tool writes to the bucket, how does the mount find out?
A vendor who answers all five without hedging has built the metadata service, the cache and the consistency rules. One who answers only the first has built an adapter, which is fine as long as it is described that way. The glossary defines the terms; the AI data pipelines guide covers the data before it reaches the bucket.
Runix FS is one implementation of this layer: metadata replicated with Raft, a cache in memory, SSD and HDD tiers on the workers, and two mount modes that make the consistency rule explicit, with the bucket as the source of truth or with operations completing in the cluster and syncing afterwards. It is built on the Apache 2.0 Curvine project and is in early access, deployed per customer in their own cloud account; the quick start brings up a single-node cluster and mounts a bucket.
Questions this raises
Is S3 a file system?
No. S3 is a key-value store for objects: it has no directories, no rename and no partial writes. Those exist only when a layer in front of it provides them.
What does POSIX mean for a file system built on object storage?
It means the layer implements the operations programs expect, such as open, seek, rename and readdir, with a metadata service and a cache behind them. Full compliance is rare, so ask which operations are emulated and which are absent.
Can other tools keep reading the bucket once a file system layer is in front of it?
Usually yes, if paths map one-to-one to object keys and the bucket keeps its layout. Whether those tools see recent writes depends on whether the mount writes through to the bucket or syncs to it later.