Docs

How outages are detected

One failed check never wakes anyone up. A dropped packet in one city is not an outage, and a false SMS at 3 a.m. is the most expensive mistake a monitoring tool can make.

Locations

We check from Frankfurt, New York, San Francisco. Each monitor has a primary location, which checks on your interval, and, depending on your plan, secondary locations, which check less often. Secondary locations exist to confirm an outage from more than one place and to compare speed around the world, not to detect it.

From suspicion to incident

  1. Suspicion. The primary location sees a failure (timeout, connection refused, 5xx, a missing keyword, a suspended hosting or parking page…). Nothing is sent yet; the incident is only investigating.
  2. Verification. Right away, not on the next scheduled check, up to two other healthy locations of the monitor check the same target.
  3. Confirmed. If most of the locations that can vote agree, the incident is confirmed and your channels are alerted. A monitor with only one location needs two failures in a row.
  4. Resolved. When checks pass again, the incident is resolved and the same channels hear about it.

What does not count

  • Our own problems. If one of our locations stops reporting, its results do not count, and it never opens an incident on your monitors.
  • Locations that never reached your target. A location that has never succeeded for a monitor does not vote on it. The cause is often a regional firewall.
  • Firewalls. A block by Cloudflare or another firewall is a warning, not an outage. See allowlisting.
  • Passwords. A site that rejects the password we send is a warning, not an outage. A site that was open and now asks for a password is an outage. See password-protected sites.
  • Maintenance windows. Checks keep running, alerts do not, and the time is left out of uptime. See maintenance windows.

Writing down the cause

On the Incidents tab of a monitor you can add a cause to a confirmed incident, for example "outage at the hosting provider, fixed on their side". It stays next to the incident and, unless you mark it as internal, appears next to the outage in the client's monthly report and in the client portal. Adding or changing a cause sends no alert and does not change the incident.

Checking on demand

After you change a monitor's settings you do not have to wait for the next scheduled check: Check now at the top of the monitor page runs one check from the primary location and shows the result there within a few seconds. You can run up to 5 in 10 minutes per monitor. A manual check is a normal check: it appears in the log, and a failure starts the same verification as above, so one failed manual check does not alert anyone on its own.

Never connected

A target that has never answered at all (usually a typo in the hostname, or a firewall in front of a TCP port) is shown as never connected. You get one notice, no incident, and uptime is not counted until the first successful check. After the first hour it is checked every 5 minutes from the primary location only. Use Test now on the monitor page after you fix it.

Uptime

Uptime is the time without a confirmed outage. Single failed checks, warnings, firewall blocks, rejected passwords and maintenance windows do not lower it. The share of successful checks per location is shown separately as Checks OK.