how to set up synthetic checks for critical user flows
Builds synthetic monitoring that catches outages before users report them. Use when you need uptime proof for login, checkout, or signup flows, when SLIs need an outside-in signal, or when status pages need real data. Covers what to script, where to run from, and alerting. Not for load testing or RUM.
TL;DR
Synthetic checks are scripted browsers or API calls that run your critical user flows every few minutes from outside your infrastructure. They catch the outages your internal metrics miss: DNS failures, CDN misconfigurations, broken third-party scripts. Start with login and checkout, run from two or more locations, and alert on consecutive failures, not single ones.
The query
how to set up synthetic checks for critical user flowsUse this when
- You need to know about outages before users tell you
- SLIs need an outside-in measurement to pair with server metrics
- A status page needs automated, trustworthy uptime data
- Deploys should be validated against real user paths
Not for when
- Load or stress testing (different tool, different goal)
- Real user monitoring (complements synthetics, does not replace them)
- Debugging a specific failing endpoint
Steps
Step 1: Pick the flows that equal "the site is down"
List the flows whose failure means an incident: usually login, signup, search, add-to-cart, checkout. If it is not on this list, a synthetic check for it can wait. Two to five flows is the right starting set. Expected output: a short list of flows, each with a definition of success (e.g. "lands on the order confirmation page").
Step 2: Script each flow as simply as possible
Write the check to do the minimum that proves the flow works: for an API flow, a login call plus one authenticated request. For a browser flow, load the page, fill the form, submit, assert on the result. Keep credentials in the synthetic tool's secret store, never in the script. Expected output: a script that passes reliably for a week with zero changes. Flaky synthetics train everyone to ignore them.
Step 3: Run from at least two locations
Run each check from two or more geographic locations or networks. One location failing is a network problem; all locations failing is your outage. This distinction saves a lot of 3am confusion. Expected output: per-location results visible separately, so regional issues are obvious.
Step 4: Alert on consecutive failures, not single blips
Alert when a check fails 2-3 times in a row, not on the first failure. Single failures are usually the synthetic infrastructure hiccuping. Tune the run frequency (every 2-5 minutes is typical) so consecutive-failure alerts still fire within your SLO detection budget. Expected output: pages that correlate with real incidents, not noise.
Step 5: Wire results into deploys and status
Show synthetic results on deploy dashboards so a bad deploy is visible within minutes, and feed uptime into the status page automatically. Synthetics that nobody looks at are just spend. Expected output: deploys that break a critical flow get caught by the synthetic check before users report it.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_qrZGdDDYm9HhdlPHlkqrOg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.