What it does
Watchtower checks every connected site's availability on a fixed schedule and confirms an outage with a second check before anyone is paged, so a blip that recovers before you'd have opened the alert never sends one. A site's Uptime tab shows stat tiles for the selected period — uptime %, total downtime, incident count, MTTR (mean time to recovery), longest incident, and average/p95 response time — plus a response-time chart with incident windows overlaid, and an incident list you can export.
The portfolio-wide Uptime page (in the main nav) rolls all of that up: fleet-wide uptime %, sites with open incidents, and the worst sites by downtime for the period.
How to
- Open a site's Uptime tab and pick a period: 24h, 7d, 30d, 90d, 12mo, or a custom range — your choice is saved in the URL.
- Set an SLA target (default 99.9%) — the chart shows a met/missed badge against it, the same number worth pasting into a client report.
- Filter and export the incident list by region or cause.
Limits & defaults
| Setting | Default |
|---|---|
| Check regions | One region today; the probe agent is built to run in more, and every result carries its region so per-region breakdowns work as regions are added |
| Confirmation delay | An incident is confirmed and re-checked before an email goes out, so a recovery inside that window sends nothing |
| Minimum outage to alert on | Very short blips are recorded but never emailed |
| Alert storm control | At most one "went down" email per site per window even during a flapping outage; the recovery email says how many were held back |
| Default SLA target | 99.9% |
| Raw check retention | 30 days; hourly/daily rollups kept 2 years for longer periods |