how to write a detection rule for brute-force logins
How to write a brute-force login detection rule: picking the log source, defining the signal as failed attempts per user or IP over a time window, baselining thresholds against historical logs, and wiring the alert with suppression. Use when password spraying shows up in logs, you are standing up detection for a new service, or an audit asks for account-takeover coverage. Triggers: 'brute force detection', 'failed login alert'. Not for: blocking traffic or tuning impossible-travel alerts.
how to write a detection rule for brute-force logins
TL;DR
A brute-force detection rule is just counting: failed logins per user or per IP over a time window, alert when the count crosses a threshold you chose from real data. Baseline first by counting historical failures so your threshold sits above background internet noise but below real attacks. Write two rules, one for brute force (many attempts, one account) and one for spraying (one source, many accounts), and suppress repeats so one attack is one alert, not five hundred.
how to write a detection rule for brute-force loginsUse this when
- Password spraying or credential stuffing shows up in your logs
- You are standing up detection for a new service or IdP
- An audit asks for account-takeover detection coverage
- Users report account lockouts they did not cause
Not for this skill when
- You want to block the traffic automatically (that is the rate-limiter and WAF skill)
- You are tuning impossible-travel or other behavioral alerts
- You need full SIEM engineering (this is the rule logic, not the platform setup)
Steps
1. Confirm you have the log data
grep -c "Failed password" /var/log/auth.logExpected: a nonzero count. Zero means the log source is wrong, rotated, or the service logs elsewhere. For cloud IdPs, confirm sign-in logs are flowing to wherever your rules run before writing a single threshold.
2. Baseline the background noise
grep "Failed password" /var/log/auth.log | awk '{print $(NF-3)}' | sort | uniq -c | sort -rn | head -20Expected: the top offending IPs with their failure counts. Internet-facing SSH gets constant background guessing, so your threshold must sit above that noise. If the top IP has 40 failures a day of random usernames, a threshold of 10 per 5 minutes still catches real targeted attacks without paging anyone for the background hum.
3. Define the two signals
Brute force is many attempts against one account. Spraying is one source trying many accounts. Catch both:
awk '/Failed password/ {ip=$(NF-3); user=$9; print ip, user}' /var/log/auth.log | sort | uniq -c | sort -rn | head -20Expected: IP and user pairs ranked by failure count. Pairs with counts well above your step 2 baseline are your candidate alerts. In SIEM terms: count failed logins grouped by target user over 5 minutes with threshold 10, and count distinct target users grouped by source IP over 10 minutes with threshold 10. Tune the numbers from your step 2 baseline.
4. Add suppression so one attack is one alert
grep "Failed password" /var/log/auth.log | awk '{print $(NF-3)}' | sort -u | wc -lExpected: the distinct attacker IP count. If a handful of IPs produce most failures, suppress per IP per hour rather than alerting per event. A rule that fires 500 times for one attack gets muted by humans, which is worse than no rule.
5. Define what happens on alert
Write the runbook before the rule goes live: who gets paged, what counts as confirmed (a successful login from the same IP right after the failures is the big one), and the response steps (force a password reset, revoke sessions, block the IP at the edge). A detection with no response plan is just expensive logging.
Variant: cloud IdP version
Query sign-in logs for failed authentication error codes grouped by user and IP over the same windows. Same two-signal structure, the field names change.
Variant: web app version
Count 401 and 403 responses on your login endpoint per IP in the access logs. This catches credential stuffing against the app even when the IdP is not involved.
Variant: fail2ban as the automated companion
Detection tells you, fail2ban acts: it watches the same logs and temporarily bans offending IPs. Pair the alert rule with it so the obvious attacks get blocked while the interesting ones still reach a human.
Why this happens
Passwords are still everywhere, and guessing them at scale is cheap. Attackers run through leaked credential lists against every login surface they can find, and the only signal you get is a pile of failed attempts that looks a lot like background noise until you count it properly. The threshold is the whole game: too low and the rule is unactionable, too high and real attacks walk under it.
Edge cases and pitfalls
- NAT and corporate egress IPs make many users look like one attacker. Allowlist or separately threshold your own egress ranges.
- Service accounts with expired passwords trip brute-force rules constantly. Exclude them or fix the password.
- Spraying and brute force need different responses. Locking the targeted account helps against brute force and does nothing against spraying.
- Attackers slow down to dodge thresholds. Keep a weekly low-and-slow review alongside the real-time rule.
- Successful logins mixed into the failure window are the highest-priority signal. Alert on those immediately, not just the failures.
- Log timestamps in mixed timezones will break your windows. Normalize to UTC in the pipeline.
Tool notes: field positions in the examples are for OpenSSH auth.log. Adapt to your IdP or app log schema.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstEAyFafX2t3QHLdPwNKb-Q
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.