pandas resample irregular timestamps fill gaps
Resamples irregular timestamp series in pandas and fills the gaps correctly. Use when resample produces NaN buckets for missing periods, when you need forward-fill vs interpolation semantics, or when irregular event data must become a regular series. Not for asof joins between series, for timezone-aware vs naive comparison errors, or for exploding list columns.
TL;DR
Set the timestamp as a DatetimeIndex, call .resample("1h") (or your frequency), then choose the fill deliberately: .ffill() for state-like data, .interpolate() for measurements, .asfreq() with no fill when gaps must stay visible.
pandas resample irregular timestamps fill gapsUse this when
- Irregular event data needs to become a regular time series
- Resample buckets come back NaN and you need a fill policy
- You are unsure whether to forward-fill, interpolate, or leave gaps
Not for this skill when
- You need nearest-timestamp matching between two series, thats merge_asof
- Timezone-aware and naive datetimes fail to compare
- You need one row per list element
Steps
- Put the timestamps on the index and sort. Resample requires a DatetimeIndex:
df["ts"] = pd.to_datetime(df["ts"])
df = df.set_index("ts").sort_index()
print(df.index.dtype, df.index.is_monotonic_increasing)Expected output: datetime64[ns] True. If to_datetime fails on some rows, you have mixed formats, coerce and inspect the NaT rows first.
- Resample to your target frequency and aggregate:
hourly = df.resample("1h")["value"].mean()
print(hourly.head(10))
print("NaN buckets:", hourly.isna().sum())Expected output: one row per hour with means where data existed and NaN where nothing was recorded. The NaN count tells you how gappy the series is.
- Pick the fill that matches the data semantics. For state-like values (a gauge, a status), forward-fill:
filled = hourly.ffill()Expected output: each NaN bucket takes the last known value. Correct for "the temperature was still X" but wrong for counters.
- For measurements between known points, interpolate instead:
filled = hourly.interpolate(method="time")Expected output: NaN buckets get time-weighted values between neighbors. Use limit= to cap how many consecutive buckets may be interpolated, so a week-long outage does not get paved over.
- When gaps are meaningful (missing data is data), keep them and flag them:
hourly = df.resample("1h")["value"].mean()
flagged = hourly.to_frame("value").assign(was_missing=lambda d: d["value"].isna())Expected output: the series plus a boolean column marking imputed vs observed buckets. Downstream aggregations can then exclude or downweight the imputed ones.
Variant phrasings
pandas resample missing dates fill
Same fix: resample creates the full date range, then .ffill(), .bfill(), or .interpolate() fills it. Choose by semantics (steps 3-5).
pandas resample upsampling NaN
Upsampling (e.g. daily to hourly) always creates NaN buckets by construction. That is not an error, it is the fill step doing its job.
irregular time series to regular pandas
Set the DatetimeIndex and resample (steps 1-2). For event counts rather than values, use .size() instead of .mean().
Why it happens
Resample bins the index into fixed-frequency buckets and aggregates per bucket. Buckets with no observations aggregate to NaN, which is honest: pandas does not know what happened in that hour. The fill method is a modeling choice about the unobserved periods, and different data types demand different choices.
Edge cases
- Daylight-saving transitions create ambiguous or nonexistent local times, resample in UTC and convert for display.
ffillbefore the first observation does nothing, leading NaNs needbfillor must stay missing.- Interpolating across huge gaps fabricates data, always set
limiton production pipelines. - Duplicate timestamps aggregate fine in resample, but check they are genuine duplicates and not a join fanout.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_Ouu7fZ8UbFOB8qg8db0h9Q
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.