## TL;DR
When the Datadog agent stops reporting, check in this order: is the process running, is the config valid, can it reach the Datadog intake, and is the API key present and correct. Nine times out of ten it is a bad config push or an expired key, not a network problem. Do not reinstall first; reinstalling wipes the diagnostics you need.

## Error / query
```text
Datadog agent not reporting: troubleshooting
```

## Use this skill when
- Hosts or containers vanish from the Datadog infrastructure list
- Metrics stop flowing but the app itself is healthy
- You just changed datadog.yaml and the agent went silent
- A new host never showed up in Datadog after install

## Not for this skill when
- The agent reports but specific integrations show no data; that is an integration config issue
- You are troubleshooting the Datadog web app or API; this is agent-side only
- APM traces are missing but infra metrics flow; check the trace agent port separately

## Steps

### Step 1: Check whether the agent process is running
```bash
systemctl is-active datadog-agent || service datadog-agent status
```
Expected: "active"; if it says inactive or failed, start it and watch the logs before doing anything else.

### Step 2: Run the agent's own status check
```bash
datadog-agent status 2>&1 | head -40
```
Expected: a component list with most showing green/OK; note any component marked with an error, especially the forwarder.

### Step 3: Validate the config file syntax
```bash
datadog-agent configcheck 2>&1 | head -30
```
Expected: "All providers and files valid" or similar; a YAML error here explains a silent agent after a config change.

### Step 4: Confirm an API key is configured (without printing it)
```bash
grep -c 'api_key' /etc/datadog-agent/datadog.yaml
```
Expected: a count of 1 or more; 0 means the key is missing entirely, which is the most common cause on new hosts.

### Step 5: Test connectivity to the Datadog intake
```bash
curl -s -o /dev/null -w '%{http_code}' -m 10 https://app.datadoghq.com/api/v1/validate
```
Expected: a 403 response means the network path works (the key is just not validated by that endpoint); a timeout means a firewall or proxy is blocking the agent.

### Step 6: Read the last agent errors
```bash
grep -iE 'error|forwarder.*fail|api.*key' /var/log/datadog/agent.log | tail -10
```
Expected: concrete errors like forwarder failures or key issues; fix what the log names, not what you guess.

## Variant phrasings

### "Datadog agent running but no metrics in the UI"
Usually a forwarder or intake problem; steps 5 and 6 find it. Also check the host is not muted in the UI.

### "Datadog agent fails after datadog.yaml change"
Run configcheck first; a single bad indent silently kills metric submission while the process stays up.

### "Datadog agent not reporting in Docker"
The container needs the API key env var and access to the Docker socket; check both before touching the host agent.

## Why it happens
The agent is a forwarder with a config file: it needs a valid key, valid YAML, and a network path to the intake. Config management pushes, key rotations, and proxy changes each break exactly one of those, and the agent fails quietly rather than loudly.

## Edge cases and pitfalls
- Do not paste the real API key into chat logs, tickets, or flare output; redact it first.
- A proxy env var set for the shell but not for the agent service is a classic silent failure; set proxy in datadog.yaml instead.
- Multiple agents on one host (host install plus container) can fight over the same hostname and look like flapping.
- After fixing, allow 2-3 minutes before declaring victory; the agent backfills on its own schedule.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_dWjgmchaWb3Z1HQGzlkUPg
