The small files problem in machine learning: millions of samples, one metadata bill

Most training datasets start life as one file per sample: a JPEG per image, a JSON record per conversation, a WAV per utterance, a source file per function. That is the natural output of collection and labelling tools, and the worst possible shape for storage. The small files problem in machine learning is that the cost of a dataset is set by how many objects it has, not by how many bytes, and millions of objects cost far more than the same bytes in larger ones.

Where the small files come from

Sample-per-file is convenient for every step before training: annotators work on one item at a time, collection scripts write one output per input, deduplication wants a stable identity per record, and a reviewer can open any sample with an ordinary tool. Code datasets are the extreme case: a repository is thousands of files, most of them small, and a code dataset built from many repositories inherits that shape multiplied.

The problem appears only at the boundary where the pipeline hands the data to a loader. The pipeline stages that produce samples are per record by nature; the training loop that consumes them wants bulk, and whichever side converts between the two pays.

Why the small files problem in machine learning is about overhead, not bytes

Every file carries fixed costs that do not shrink with its size. In an object store, each one is a request: request signing, a network round trip, the request and the response, and in most pricing models a per-request charge on top of the bytes. On a file system, each one is an inode, an open, a read and a close, and across a FUSE boundary each of those crosses between kernel and user space.

For a large file those costs are noise next to the transfer; for a small one they are the transfer. A loader reading thousands of small samples per step is bound by how many requests per second the storage path can answer, not by bandwidth, and the GPU behind it waits on request latency rather than on bytes.

Listing and metadata pressure

Before a job reads its first sample it needs to know what the samples are. On an object store that means listing a prefix, which is paged and sequential: millions of keys take many round trips to enumerate, every job repeats the work, and the list is stale by the time it finishes if anything is still writing.

On a file system the listing is faster but the metadata becomes the bottleneck instead. Every file is an entry a metadata service has to hold and answer for, and a training job's opens, stats and readdirs are a storm of small metadata operations at the rate the loader runs. HDFS made this failure famous: its name node keeps the namespace in memory, so many small files exhaust the name node before the disks. Any file system with a central metadata service has the same ceiling, placed higher or lower by its design.

The problem is therefore two-sided: small files on an object store are expensive per read, and small files on a file system are expensive per entry. Each remedy below moves the cost to one side or the other.

Packing: shards, Parquet and WebDataset-style formats

Packing concatenates many samples into one large file, so that a read is sequential and a listing is short. The common forms:

Packing trades away three things. Random access to one sample by id is gone; finding it means reading the shard or keeping a side index. Shuffling moves to shard granularity plus a buffer, which is weaker than a global shuffle. And updating one sample means rewriting its shard, so a dataset under active curation is repacked on every change, and per-record identity survives only if it is carried inside the shard.

A file system with a metadata service

The alternative is to leave the samples as files and put a file system in front of the bucket that is built to absorb the metadata load. Its metadata service holds the namespace and answers stat and readdir without touching the bucket, replicated to survive a node failure. A cache on memory and SSD serves the bytes, so each small object is fetched from the bucket once, when the cache warms, and served locally on every epoch after.

This keeps what packing gives up: every sample is still addressable by path, tools that expect files keep working, a curated dataset can change one sample at a time, and the bucket keeps its one-object-per-sample layout for everything else that reads it. The per-request cost of the object store is paid once at warm time instead of once per epoch.

It does not make the metadata free. The service has a capacity in entries, a sizing question to ask before committing a dataset to it, and the cluster has to be running for the path to exist. A sequential pass over packed shards still touches fewer metadata entries per sample than any file system can, because a shard is one entry for many samples.

Choosing between packing and a file system

Whichever you choose, measure the per-sample cost on the first epoch and the second. Packing fixes the first, a cache fixes the second, and neither fixes a loader that spends its time decoding.

Runix FS is designed for the second case: a file system over an S3-compatible bucket, with metadata held by masters replicated with Raft, a cache in memory, SSD and HDD tiers on the workers, and paths that map one-to-one to object keys, so the sample-per-object layout stays readable by anything else. It is in early access, sized and deployed with each customer in their own cloud account; the quick start brings up a single-node cluster, and the introduction states what we do and do not claim.

Questions this raises

Why are small files slow to train on?

Each file carries a fixed cost, a request on an object store or a metadata entry and an open on a file system, that does not shrink with its size. With millions of small samples those fixed costs, not the bytes, set the read rate.

Should I pack my dataset into tar shards or Parquet?

Pack when the dataset is frozen and read sequentially many times by a loader you control. Tar shards suit mixed media read in order; Parquet suits tabular and text data where a reader benefits from skipping columns.

Does a file system with a metadata service solve the small files problem?

It moves the cost rather than removing it. The per-object cost of the bucket is paid once when the cache warms instead of every epoch, while the metadata service has a capacity in entries that needs sizing up front.

Related to this post: Runix FS. Tell us what you are building and we reply within one business day.