## TL;DR
Print the actual schema with `df.printSchema()` and match the column reference exactly: Spark is case-sensitive by default for DataFrame ops, nested fields need dot or backtick notation, and after a join ambiguous columns must be qualified with their table alias.

```text
org.apache.spark.sql.AnalysisException: cannot resolve '[col]' given input columns: [...]
```

## Use this when
- Spark throws AnalysisException about an unresolvable column
- A column exists in the source table but the query cannot see it
- Struct fields or post-join columns fail to resolve

## Not for this skill when
- Executors run out of memory during a shuffle
- `to_timestamp` returns nulls on bad rows
- You are tuning broadcast join thresholds

## Steps

1. Print the real schema. The error message lists input columns, but the schema shows nesting and exact case:

```python
df.printSchema()
```
Expected output: the full column tree with exact names and types. Compare character by character with the name in your query, most failures are a case or underscore difference.

2. Fix the reference. The three usual culprits:

```python
# case: Spark DataFrame API is case-sensitive even when SQL is not
# nested struct field:
df.select("address.city")
# or with backticks when names contain dots or spaces:
df.select("`user.name`")
```
Expected output: the select succeeds. For struct fields, `parent.child` works; for columns whose names literally contain dots, backticks are required.

3. After a join, qualify ambiguous columns with the side they came from:

```python
orders.alias("o").join(customers.alias("c"), "customer_id") \
    .select("o.order_id", "c.name")
```
Expected output: no ambiguity error. The bare column name existed on both sides, so Spark refused to guess.

4. If the column truly is missing, find where it got dropped:

```python
for stage, frame in [("raw", raw), ("cleaned", cleaned), ("joined", joined)]:
    print(stage, frame.columns)
```
Expected output: the column lists per pipeline stage, showing exactly which transformation dropped or renamed it. A `select` upstream is the usual suspect.

5. For SQL strings, remember quoting rules differ from the DataFrame API:

```sql
SELECT `weird column`, address.city FROM events
```
Expected output: backtick-quoted identifiers resolve. In Spark SQL, double quotes are string literals by default, not identifier quotes.

## Variant phrasings

### spark cannot resolve column after withColumnRenamed
The rename did not apply where you think, or you renamed on a different frame than the one you query. Print columns right before the failing select (step 4).

### AnalysisException cannot resolve due to case
Set `spark.sql.caseSensitive=false` if you truly want case-insensitive SQL, but the DataFrame API stays case-sensitive regardless.

### spark struct field cannot resolve
Use dot notation `parent.child` for real nesting, backticks for literal dots in names. `getField("child")` is the programmatic alternative.

## Why it happens
Spark resolves column references against the logical plan's output schema at analysis time, before any data is read. Anything that changes names between plan construction and the reference, case differences, renames, drops, ambiguous joins, or struct flattening, fails here with AnalysisException instead of at runtime.

## Edge cases
- Column names with spaces or special characters need backticks in SQL and work unquoted in the DataFrame API only if they are valid identifiers.
- `df["col"]` and `df.col` behave the same for resolution, but `df.col` breaks on names that are not valid Python identifiers.
- After `explode`, the original array column is gone, reference the exploded column instead.
- Delta Lake column mapping can rename physical columns, check the table's column mapping mode when the table schema looks right but resolution fails.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_NuYpaBil0H6Oz1oEP7PR8Q
