## TL;DR
Set the timestamp as a DatetimeIndex, call `.resample("1h")` (or your frequency), then choose the fill deliberately: `.ffill()` for state-like data, `.interpolate()` for measurements, `.asfreq()` with no fill when gaps must stay visible.

```text
pandas resample irregular timestamps fill gaps
```

## Use this when
- Irregular event data needs to become a regular time series
- Resample buckets come back NaN and you need a fill policy
- You are unsure whether to forward-fill, interpolate, or leave gaps

## Not for this skill when
- You need nearest-timestamp matching between two series, thats merge_asof
- Timezone-aware and naive datetimes fail to compare
- You need one row per list element

## Steps

1. Put the timestamps on the index and sort. Resample requires a DatetimeIndex:

```python
df["ts"] = pd.to_datetime(df["ts"])
df = df.set_index("ts").sort_index()
print(df.index.dtype, df.index.is_monotonic_increasing)
```
Expected output: `datetime64[ns] True`. If `to_datetime` fails on some rows, you have mixed formats, coerce and inspect the NaT rows first.

2. Resample to your target frequency and aggregate:

```python
hourly = df.resample("1h")["value"].mean()
print(hourly.head(10))
print("NaN buckets:", hourly.isna().sum())
```
Expected output: one row per hour with means where data existed and NaN where nothing was recorded. The NaN count tells you how gappy the series is.

3. Pick the fill that matches the data semantics. For state-like values (a gauge, a status), forward-fill:

```python
filled = hourly.ffill()
```
Expected output: each NaN bucket takes the last known value. Correct for "the temperature was still X" but wrong for counters.

4. For measurements between known points, interpolate instead:

```python
filled = hourly.interpolate(method="time")
```
Expected output: NaN buckets get time-weighted values between neighbors. Use `limit=` to cap how many consecutive buckets may be interpolated, so a week-long outage does not get paved over.

5. When gaps are meaningful (missing data is data), keep them and flag them:

```python
hourly = df.resample("1h")["value"].mean()
flagged = hourly.to_frame("value").assign(was_missing=lambda d: d["value"].isna())
```
Expected output: the series plus a boolean column marking imputed vs observed buckets. Downstream aggregations can then exclude or downweight the imputed ones.

## Variant phrasings

### pandas resample missing dates fill
Same fix: resample creates the full date range, then `.ffill()`, `.bfill()`, or `.interpolate()` fills it. Choose by semantics (steps 3-5).

### pandas resample upsampling NaN
Upsampling (e.g. daily to hourly) always creates NaN buckets by construction. That is not an error, it is the fill step doing its job.

### irregular time series to regular pandas
Set the DatetimeIndex and resample (steps 1-2). For event counts rather than values, use `.size()` instead of `.mean()`.

## Why it happens
Resample bins the index into fixed-frequency buckets and aggregates per bucket. Buckets with no observations aggregate to NaN, which is honest: pandas does not know what happened in that hour. The fill method is a modeling choice about the unobserved periods, and different data types demand different choices.

## Edge cases
- Daylight-saving transitions create ambiguous or nonexistent local times, resample in UTC and convert for display.
- `ffill` before the first observation does nothing, leading NaNs need `bfill` or must stay missing.
- Interpolating across huge gaps fabricates data, always set `limit` on production pipelines.
- Duplicate timestamps aggregate fine in resample, but check they are genuine duplicates and not a join fanout.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_Ouu7fZ8UbFOB8qg8db0h9Q
