There are three ways to get a bucket in front of a training job or an agent on Kubernetes: mount it with FUSE, provision it through a CSI driver, or skip the mount and call an SDK or a Hadoop-compatible client from the code. Each changes a different thing, your code, your pod spec or your operations, and each fails in its own way. This post is about how to mount object storage in Kubernetes for ML work, and when not to mount it at all.
Three ways to mount object storage in Kubernetes
The three are not alternatives at the same level. FUSE is a mechanism: a kernel interface that lets a user-space process answer file system calls. CSI is a Kubernetes contract for provisioning storage and attaching it to pods, and a CSI driver often uses FUSE underneath. An SDK is a library, with no kernel or cluster involvement.
So the question of which one is really three questions: who manages the mount, what the code sees, and where failures surface.
FUSE: a path, at the price of a daemon
With a FUSE mount, a user-space process registers with the kernel through libfuse or an equivalent, and every open, read, readdir and stat on the mount point is forwarded to that process, which answers from a cache or from the bucket. The code does not change: a loader that opens a path keeps opening a path, and shell tools, notebooks and libraries all work.
What you pay for that is the daemon. Every file operation crosses from the kernel to user space and back, a fixed cost per call that small files and metadata-heavy workloads feel most. The daemon runs on every node that mounts, needs access to the FUSE device, and holds both the cache and the credentials.
The failure modes are specific. If the daemon dies, processes with open files see errors such as "transport endpoint is not connected", and the mount point has to be remounted, which Kubernetes does not do on its own. If the bucket is slow, a read blocks inside the kernel and often cannot be interrupted cleanly. And a daemon that caches aggressively competes with the training job for the node's memory, in exactly the place you cannot afford contention.
CSI: the same mount, owned by the cluster
The Container Storage Interface is the contract that lets a storage system provision volumes and mount them into pods without changes to Kubernetes itself. A driver for a file system over object storage typically runs a node plugin on each node, which performs the mount, and a controller that handles provisioning. The pod spec declares a PersistentVolumeClaim; the code, again, sees a path.
What changes is ownership and scale. The platform team installs and upgrades the driver, application teams request volumes in YAML, and a ReadWriteMany volume lets many pods on many nodes share one namespace. For workloads that start thousands of pods, the provisioning path matters: a driver that creates a volume by calling a cloud API inherits that API's rate limits and, for block volumes, the per-node attachment limits, while one that provisions by creating a directory in an existing file system does not.
The failure modes move with ownership. A driver upgrade can require draining nodes or remounting every volume, and a node plugin that runs the FUSE processes in its own pod takes every mount on that node with it when it crashes, so the blast radius is the node rather than the process. Provisioning or mounting that is slow at scale shows up as pods stuck in Pending or ContainerCreating with no error from the application, which is hard to alert on. And since most such drivers mount with FUSE, every FUSE failure mode is inherited.
SDK or Hadoop-compatible client: no mount, your code changes
The third path drops the mount. A Python, Java or Rust SDK, or a Hadoop-compatible FileSystem implementation for Spark and Flink, lets the application open, read and list through library calls, with the client handling caching, parallelism and retries inside the process.
Nothing runs in the kernel or as a daemon, so there is no privileged pod, no remount after a crash and no competition for node memory outside the process's own. Parallel reads are explicit rather than hidden behind a blocking read(), and errors arrive as exceptions the application can handle, with timeouts it chose. For jobs already on Spark, the Hadoop FileSystem API is the interface they expect anyway.
The cost is that every tool which expects a path stops working: image decoders, tokenisers, checkpoint writers, notebooks and the long tail of scripts all assume open(). Each team carries its own client version, so an upgrade is a code change in every repository rather than a driver roll-out, and retries and timeouts become the application's responsibility, which each team gets wrong in its own way.
The three paths side by side
| Path | What changes | Who owns it | Failure surface |
|---|---|---|---|
| FUSE mount | Nothing in the code | Whoever runs the daemon | Daemon death hangs the mount; node memory contention |
| CSI driver | The pod spec | The platform team | Node plugin crash can take the node's mounts; slow provisioning at scale |
| SDK or Hadoop client | Every file access | Each application team | Version drift; retries and timeouts per application |
When to pick which
- Pick a FUSE mount when the code cannot change and the mount is run by the same people who run the job: a research team with its own nodes, a notebook host, a single training cluster.
- Pick a CSI driver when more than one team shares a cluster, when pods need a volume declared in YAML, and when the platform team will own the driver's lifecycle. Check how provisioning scales first.
- Pick an SDK or Hadoop client when the workload is already a Spark or Flink job, when the loader is purpose-built and the team controls every read, or when a daemon with device access on the node is not acceptable.
- Pick more than one when the same data serves different consumers: a CSI volume for the training pods and the Hadoop client for the preprocessing job are a normal pairing, provided both see the same namespace.
The deciding question for all three applies to every file system layer over a bucket: where does the truth live, and what does a delete through this path do to the bucket? That is a property of the file system behind the mount, not of the mount mechanism; the introduction to Runix FS describes how one layer defines it.
Runix FS exposes all three paths over one namespace: a FUSE mount, a Kubernetes CSI driver that provisions ReadWriteMany volumes by creating a directory rather than calling a cloud API, a Hadoop-compatible client, and SDKs for Java, Python and Rust, with paths mapping one-to-one to object keys so the bucket stays independently readable. It is in early access and deployed per customer in their own cloud account; the quick start brings up a single-node cluster, attaches a bucket and mounts it.
Questions this raises
Is a CSI driver the same as a FUSE mount?
No. CSI is the Kubernetes contract for provisioning and attaching volumes; a CSI driver for object storage usually performs the actual mount with FUSE, so it adds cluster-managed provisioning on top of the FUSE mechanism and inherits its failure modes.
Which option needs no code changes?
A FUSE mount and a CSI volume both present a path, so existing loaders and tools work unchanged. An SDK or Hadoop-compatible client replaces every file access with a library call.
What happens to a FUSE mount when the daemon crashes?
Processes with open files see transport errors and the mount point stops answering until it is remounted. Kubernetes does not remount it on its own, so the recovery path has to be planned.