A multi-tier cache for AI training data: what goes where

A multi-tier cache for AI training is a set of storage tiers, usually memory, SSD and HDD on the nodes near the GPUs, holding copies of data whose durable home is an object store. The point is to make the second read of a sample local, and the first read local for every node except the one that fetched it. This post covers what belongs on each tier, how eviction and warming should work for training, how this differs from the page cache every node already has, and when a cache does nothing useful.

Why the cache sits in front of the bucket

Object storage is the right durable home for training data: cheap, durable and shared. It is the wrong place to read from on every step, because each read is a network request with its own cost and latency, paid once per sample. A cache near the GPUs changes the arithmetic: the bucket is read once per block, and later reads are served from local or near-local media.

This is not the prompt cache on the inference side, which stores attention state and is governed by prefix matching; here the cached thing is bytes, and correctness is about staleness, whether the cache still matches the bucket.

What a multi-tier cache for AI training holds on each tier

Tiers differ in capacity, cost and random-read behaviour, and the assignment follows from that.

The write path should land on the fastest tier that has room and demote from there, so that a block written once and read soon after never touches the slower media.

Eviction for an access pattern with no locality

General-purpose caches assume temporal locality: what was read recently will be read again soon. A shuffled epoch has almost none, because every sample is read exactly once per epoch, in a random order, and again next epoch in a different order. Under that pattern, least-recently-used eviction with a cache smaller than the working set degenerates into evicting each block just before it is needed again.

Two rules follow. First, size the cache to the working set, not to a hit-rate target: a cache that holds the whole epoch hits on every read after the first, and one that holds most of it can still miss on most reads. Second, protect the working set from scans: a validation run or a checkpoint write should not push the training set out, which scan-resistant policies, pinned directories or a capacity budget per job all address.

Eviction between tiers is a separate decision from eviction out of the cache: demotion from memory to SSD keeps a block close, while eviction from HDD sends the next read back to the bucket.

Warming: the first epoch is the whole problem

A cold cache does not help the first epoch at all, and for a short fine-tuning run the first epoch may be most of the job. Warming pulls the working set into the cache before the loader asks for it, and is cheaper than it looks: one read of the dataset from the bucket in large sequential requests, instead of the loader's many small random ones.

It is worth doing when the dataset is known ahead of time, which for training it almost always is; trigger it from the job definition, not by hand, and make it idempotent so that a retried job does not fetch twice. With a shared cache, warming once serves every node in the job.

Page cache versus distributed cache

Every Linux node already has a cache: the kernel's page cache, which keeps recently read file pages in otherwise unused memory. It is free, automatic and effective for one node whose dataset fits in memory, and the wrong tool for a cluster, for four reasons.

A distributed cache fixes all four: it is shared across nodes, has explicit capacity on each tier, survives process restarts, and knows which blocks it holds and which are still only in the bucket. It costs a network hop on a hit served by another node and a metadata lookup to find the block, which is why local memory remains the first tier even in a distributed design.

When a cache is pointless

A cache earns its keep on rereads. Where there are none, it adds a hop and a bill.

The honest test is to measure the hit rate on the second epoch and the step time on the first. If the second epoch is not faster, the cache is not doing anything; if the first epoch is the whole job, warming is the feature to evaluate, not the tiers.

Runix FS is a cache of this kind: memory, SSD and HDD tiers on the workers, written to the fastest available tier first, in front of an S3-compatible bucket whose layout stays unchanged, built on the Curvine project. It is in early access, sized with each customer for their working set and deployed in their own cloud account; the quick start brings up a single-node cluster, and the introduction covers what we will and will not claim about it.

Questions this raises

What is the difference between the page cache and a distributed cache for training data?

The page cache is per node, memory only, unmanaged and unaware of the bucket. A distributed cache is shared across nodes, has explicit capacity on memory, SSD and HDD tiers, survives restarts and tracks which blocks it holds.

How big should a cache for AI training data be?

Large enough to hold the working set an epoch touches. Because shuffled training has almost no temporal locality, a cache that holds only part of the working set can miss on most reads.

When does a cache not help a training job?

When the data is read once and never again, when the loader is CPU-bound on decoding, when the working set cannot fit, or when the data already sits on local NVMe that is large enough.

Related to this post: Runix FS. Tell us what you are building and we reply within one business day.