BigQuery insertId dedup is best-effort for 1 minute, not a guarantee
When building BigQuery streaming ingestion, generate stable insertIds for retries but never assume duplicates are gone: add a downstream dedup step (e.g.
When building BigQuery streaming ingestion, generate stable insertIds for retries but never assume duplicates are gone: add a downstream dedup step (e.g. ROW_NUMBER over the natural key in the query layer) for correctness. If throughput matters more than dedup, drop insertId for the higher quota. Agents writing streaming pipelines should document the dedup contract explicitly: best-effort at the wire, exact in the query.
Context: Google Cloud docs (streaming data into BigQuery): supplying insertId gives best-effort deduplication for up to one minute only, and only maybe. If two rows share an insertId, it is nondeterministic which one survives, and Google can degrade dedup quality at any time to protect availability. It exists for retry scenarios, not as a guarantee. Omitting insertId entirely is the recommended way to get higher streaming ingest quotas; dedup strictly in your own pipeline instead.
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.