## TL;DR
Raise the limit in three places together: broker/topic `message.max.bytes`, producer `max.request.size`, and consumer `fetch.max.bytes` plus `max.partition.fetch.bytes`. The error tells you which side rejected the message, and the fix is always aligning all three.

```text
org.apache.kafka.common.errors.RecordTooLargeException: The message is 2097164 bytes when serialized which is larger than 1048576
```

## Use this when
- Producers fail with RecordTooLargeException
- Consumers fail fetching with a fetch-size error
- You are raising the max message size deliberately

## Not for this skill when
- You are setting up consumer lag monitoring
- Consumers rebalance in a storm
- You need exactly-once semantics

## Steps

1. Read which side rejected the message. Producer-side:

```text
RecordTooLargeException: The message is 2097164 bytes when serialized which is larger than 1048576
```
Expected output: the producer's `max.request.size` (default 1MB) rejected it. The fix starts here, not at the broker.

2. Raise the producer limit above your largest message:

```python
producer = Producer({
    "bootstrap.servers": "YOUR_HOST:9092",
    "max.request.size": 5242880,  # 5MB, must exceed largest single message
})
```
Expected output: the producer accepts the message. `max.request.size` caps the whole produce request, so it must exceed the largest single message plus batching overhead.

3. Raise the broker and topic limit to match, otherwise the broker rejects what the producer now sends:

```bash
kafka-configs.sh --bootstrap-server YOUR_HOST:9092 --entity-type topics \
  --entity-name events --alter --add-config message.max.bytes=5242880
```
Expected output: the topic accepts up to 5MB messages. The broker default `message.max.bytes` is 1MB; the topic override keeps the change scoped instead of cluster-wide.

4. Raise the consumer fetch limits so consumers can actually read the large messages:

```python
consumer = Consumer({
    "bootstrap.servers": "YOUR_HOST:9092",
    "fetch.max.bytes": 52428800,      # 50MB per fetch
    "max.partition.fetch.bytes": 5242880,  # 5MB per partition
})
```
Expected output: consumers fetch large messages without errors. `fetch.max.bytes` must exceed the largest message or the consumer stalls on that partition forever.

5. Check the replica fetch path too, it has its own limit:

```bash
# broker config: replica.fetch.max.bytes must be >= message.max.bytes
```
Expected output: replicas can replicate the large messages. Forgetting this one produces under-replicated partitions instead of a clear error.

## Variant phrasings

### kafka message too large error
Align all three layers (steps 2-4). Fixing only the producer moves the failure to the broker; fixing only the broker moves it to the consumer.

### kafka increase max message size
Topic-level `message.max.bytes` plus client settings. Prefer topic scope over cluster-wide broker changes.

### kafka consumer stuck on large message
`fetch.max.bytes` smaller than the message means the consumer can never fetch it and the partition stalls. Raise it above the max message size.

## Why it happens
Kafka enforces size limits independently at three layers: what the producer will send, what the broker will accept, and what the consumer will fetch. The defaults (1MB) date from an era of small events, and any layer left at default becomes the ceiling no matter what the others allow.

## Edge cases
- Compression changes the serialized size, measure with your actual compression codec enabled.
- `message.max.bytes` applies per message batch after compression on the broker side.
- Very large messages hurt throughput for everyone on the partition, consider a claim-check pattern (store the payload elsewhere, send a reference) past ~10MB.
- MirrorMaker and other replicating clients need the same fetch limits as regular consumers.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_oTQHhqZKVCEkIJCzaWemKw
