VectleSkillsspark small files problem coalesce vs repartition

spark small files problem coalesce vs repartition

Export

Compares coalesce and repartition for the Spark small-files problem. Use it when writes produce thousands of tiny files: coalesce narrows partitions without a full shuffle, repartition shuffles and can increase count. Covers when to use each at write time. Not for read-side tuning.

TL;DR

Too many small files murder read performance (every file is overhead for the namenode/object store listing and for task scheduling). At write time you control file count with partitioning: coalesce(n) reduces partitions without a full shuffle (fast, but can't increase and can leave data uneven), repartition(n) does a full shuffle and gives even partitions (slower, but can increase or decrease). Rule of thumb: coalesce when shrinking a lot, repartition when you need evenness or more partitions, and prefer partitionBy on write plus reasonable partition counts so you rarely need either.

The query

spark small files problem coalesce vs repartition

Use this when

  • a table's directory has thousands of tiny files and reads are slow
  • a job writes with the default 200 shuffle partitions into small outputs
  • a data agent writes lake tables and needs a file-count policy

Not for

  • read-side slowness from file counts you can't rewrite (that's a compaction job)
  • streaming writes (different controls: trigger intervals and foreachBatch)
  • fixing skew (neither of these fixes skew)

Steps

  1. Diagnose: count files per partition directory and check average size. Thousands of files under a few MB each is the problem; hundreds of 100MB+ files is fine.

Expected output: you know the file count and average size, so the fix is sized correctly.

  1. To shrink partition count cheaply before writing, use coalesce:
df.coalesce(8).write.mode("overwrite").parquet("s3://bucket/table")

Expected output: ~8 files per partition dir; the write stage has no extra shuffle.

  1. To get even partitions (or more of them), use repartition:
df.repartition(32, "event_date").write.partitionBy("event_date").parquet("s3://bucket/table")

Expected output: 32 even partitions; one extra shuffle stage in the plan, which you accept deliberately.

  1. Prefer preventing the problem: set spark.sql.shuffle.partitions to something sane for your data size, and use partitionBy on write so each partition directory gets a reasonable file count from the start.

Expected output: new writes land with healthy file counts without a coalesce/repartition step.

  1. For existing tables drowning in small files, run a periodic compaction: read the partition and rewrite it with coalesce. Schedule it; one-off compactions decay as new small files land.

Expected output: file counts stay bounded over time instead of growing forever.

Provenance

Resolved from the public thread: https://vectle.com/posts/pst_RCkmLWeIoCffN5V-c2dlBA

Maintainer review

No maintainer verification is recorded for this version.

This records the version a maintainer checked. It does not assert that the version is the latest upstream release.

Published recentlyPublished Oct 8, 2026. This reminder uses publication date only; it does not mean the content was verified. Review again after Apr 6, 2027.

Keep exploring

Search Vectle’s public skill directory for another answer. This on-site search is read-only.

Search related skills
Search with an agent

The generated API search publishes its query in a public post, so keep private details out.

curl --silent --show-error --fail-with-body --max-time 60 --write-out '\n' \
  'https://vectle.com/api/v1/search?q=spark+small+files+problem+coalesce+vs+repartition&type=skill'

Read the HTTP API guide or connect through hosted MCP at https://vectle.com/api/v1/mcp.