chrony time sync "clock skew detected" on cloud VMs: how to fix
Fixes clock skew on cloud VMs using chrony. Use when logs show clock skew, when Kerberos or TLS fails from time drift, or when chrony cannot sync. Not for hardware clock issues.
TL;DR
Clock skew breaks time-sensitive protocols (Kerberos, TLS, database replication) long before humans notice the clock is wrong. On cloud VMs, fix chrony by pointing it at reliable sources (the cloud provider's internal NTP, plus public fallbacks), verifying it can reach them through the firewall, and letting it slew or step the clock into sync. Then monitor the offset so drift pages before it breaks auth.
The query
chrony time sync "clock skew detected" on cloud VMs: how to fixUse this when
- Applications fail with clock skew errors
- Kerberos or TLS breaks from time drift
- Chrony shows unsynchronized or large offset
- After VM snapshots or migrations
Not for when
- Hardware RTC issues on physical hosts
- Timezone configuration (different from skew)
- Application-level timestamp bugs
Steps
Step 1: Check the current offset and sync state
Query chrony's tracking info for the system offset and whether it considers itself synchronized. An offset of seconds is a nuisance; minutes break Kerberos; hours break everything. Expected output: the offset quantified and sync state known.
Step 2: Verify NTP reachability
Confirm the VM can reach its configured NTP sources: cloud metadata NTP endpoints and any public servers, on UDP 123. Egress firewalls blocking NTP are a common cloud misconfiguration. Expected output: NTP reachable, or the firewall block found.
Step 3: Use the cloud provider's internal NTP
Point chrony at the provider's internal time service first (low latency, no egress dependency), with public pools as fallback. Internal NTP is more reliable and does not depend on internet egress. Expected output: chrony synced to a low-jitter local source.
Step 4: Correct large offsets deliberately
For large offsets, decide between stepping (instant jump, can confuse running apps) and slewing (gradual, slower). Big jumps on database or Kerberos hosts deserve a maintenance window; small ones can slew. Expected output: the clock correct with the disruption understood.
Step 5: Monitor offset continuously
Alert on NTP offset exceeding a threshold well below what breaks your protocols. Clock drift is gradual; the monitoring should catch it days before Kerberos notices. Expected output: drift paging before it causes failures.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_aqgGZ0yu45GBS3U52SbC8g
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.