Object storage is not a file system, and what POSIX on object storage adds

Object storage and file systems get used interchangeably in architecture diagrams, and the confusion becomes expensive once a training job or an agent platform reads from a bucket through code written for files. This post defines both, lists what a bucket cannot do, explains what a POSIX file system on object storage has to add, and is honest about what it still cannot promise.

What object storage actually is

An object store is a key-value service for large values. You put an object under a key, get it back by that key, list the keys that share a prefix, and delete it. The key is a string; the slashes in it are a naming convention, not a directory tree.

Each object is written whole: a put replaces the entire value under that key, and there is no call that changes bytes in the middle of an existing object. Multipart upload assembles a large object from parts, but the result is still one immutable value until the next put replaces it.

In exchange you get durability, capacity that grows without planning, and a price per gigabyte a file server cannot match, which is why AI data lands in buckets and why a file system layer in front of the bucket exists as a category at all.

What a file system promises that a bucket does not

The POSIX interface, specified by the Open Group and IEEE, is what almost every data loader and shell tool was written against. Its guarantees are small and specific, and each one is missing from an object API.

What a POSIX file system on object storage has to add

A layer that presents a bucket as a file system cannot change what the bucket does; it has to add three components of its own and keep them in agreement.

A metadata service. Something has to hold the directory tree, sizes, timestamps, permissions and the mapping from a path to the blocks or objects that hold its bytes. Once that service exists, rename and directory listing become operations on it rather than on the bucket, which is where the speed comes from. It also has to survive a node failure, which is why it is usually replicated with a consensus protocol such as Raft.

A cache. Every open, read and stat that reaches the bucket pays a network round trip. A data cache on local memory and disk turns repeated reads into local reads, and a write buffer turns small writes into whole objects, which is what makes partial writes possible at all.

Consistency rules. This is the part vendors describe least and buyers should ask about most. The layer has to say where the truth lives: either the bucket is the source of truth and every write passes through to it before it is acknowledged, or the cluster is the source of truth and the bucket is updated afterwards. The second is faster, and it means the bucket is briefly, or indefinitely, behind, which decides what another tool reading the bucket directly sees.

What it still cannot do

A file system layer removes the cost of the missing operations; it does not remove their nature. Keep these limits in mind when a product page says POSIX without qualification.

Questions to ask about any POSIX layer

The claim of POSIX on S3 covers everything from a thin FUSE adapter that turns each read into a GET to a distributed file system with its own metadata and cache. The way to tell them apart is to ask five questions.

  1. Where does metadata live, and what happens to it if that node dies?
  2. What does rename do, in the cluster and in the bucket?
  3. When is a write durable, and in which place?
  4. What does a delete through the mount do to the bucket?
  5. If another tool writes to the bucket, how does the mount find out?

A vendor who answers all five without hedging has built the metadata service, the cache and the consistency rules. One who answers only the first has built an adapter, which is fine as long as it is described that way. The glossary defines the terms; the AI data pipelines guide covers the data before it reaches the bucket.

Runix FS is one implementation of this layer: metadata replicated with Raft, a cache in memory, SSD and HDD tiers on the workers, and two mount modes that make the consistency rule explicit, with the bucket as the source of truth or with operations completing in the cluster and syncing afterwards. It is built on the Apache 2.0 Curvine project and is in early access, deployed per customer in their own cloud account; the quick start brings up a single-node cluster and mounts a bucket.

Questions this raises

Is S3 a file system?

No. S3 is a key-value store for objects: it has no directories, no rename and no partial writes. Those exist only when a layer in front of it provides them.

What does POSIX mean for a file system built on object storage?

It means the layer implements the operations programs expect, such as open, seek, rename and readdir, with a metadata service and a cache behind them. Full compliance is rare, so ask which operations are emulated and which are absent.

Can other tools keep reading the bucket once a file system layer is in front of it?

Usually yes, if paths map one-to-one to object keys and the bucket keeps its layout. Whether those tools see recent writes depends on whether the mount writes through to the bucket or syncs to it later.

Related to this post: Runix FS. Tell us what you are building and we reply within one business day.