spark small files problem coalesce vs repartition
Compares coalesce and repartition for the Spark small-files problem. Use it when writes produce thousands of tiny files: coalesce narrows partitions without a full shuffle, repartition shuffles and can increase count. Covers when to use each at write time. Not for read-side tuning.
TL;DR
Too many small files murder read performance (every file is overhead for the namenode/object store listing and for task scheduling). At write time you control file count with partitioning: coalesce(n) reduces partitions without a full shuffle (fast, but can't increase and can leave data uneven), repartition(n) does a full shuffle and gives even partitions (slower, but can increase or decrease). Rule of thumb: coalesce when shrinking a lot, repartition when you need evenness or more partitions, and prefer partitionBy on write plus reasonable partition counts so you rarely need either.
The query
spark small files problem coalesce vs repartitionUse this when
- a table's directory has thousands of tiny files and reads are slow
- a job writes with the default 200 shuffle partitions into small outputs
- a data agent writes lake tables and needs a file-count policy
Not for
- read-side slowness from file counts you can't rewrite (that's a compaction job)
- streaming writes (different controls: trigger intervals and foreachBatch)
- fixing skew (neither of these fixes skew)
Steps
- Diagnose: count files per partition directory and check average size. Thousands of files under a few MB each is the problem; hundreds of 100MB+ files is fine.
Expected output: you know the file count and average size, so the fix is sized correctly.
- To shrink partition count cheaply before writing, use
coalesce:
df.coalesce(8).write.mode("overwrite").parquet("s3://bucket/table")Expected output: ~8 files per partition dir; the write stage has no extra shuffle.
- To get even partitions (or more of them), use
repartition:
df.repartition(32, "event_date").write.partitionBy("event_date").parquet("s3://bucket/table")Expected output: 32 even partitions; one extra shuffle stage in the plan, which you accept deliberately.
- Prefer preventing the problem: set
spark.sql.shuffle.partitionsto something sane for your data size, and usepartitionByon write so each partition directory gets a reasonable file count from the start.
Expected output: new writes land with healthy file counts without a coalesce/repartition step.
- For existing tables drowning in small files, run a periodic compaction: read the partition and rewrite it with coalesce. Schedule it; one-off compactions decay as new small files land.
Expected output: file counts stay bounded over time instead of growing forever.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_RCkmLWeIoCffN5V-c2dlBA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.