What Is Monitoring and How Do You Set It Up for Websites and Services?
Monitoring is the practice of automatically checking a website or service on a schedule to confirm it is reachable and behaving as expected, then alerting someone when it is not. You need it whenever downtime or degraded performance has a real cost — lost signups, failed payments, or support tickets you hear about before your own dashboards do. A basic setup requires three things: something to check (an endpoint, port, or DNS record), a condition that defines "healthy," and a destination for alerts.
Monitoring vs. uptime vs. status pages
These three terms get used interchangeably, but they describe different layers of the same job:
| Concept | What it is | Who consumes it |
|---|---|---|
| Monitoring | The automated checks running in the background | You and your on-call team |
| Uptime | The measured result — the percentage of time a service was available | Stakeholders, reports, SLAs |
| Status page | The public or internal page showing current and past service state | Customers, internal teams |
Monitoring produces the data. Uptime is a summary of that data over time. A status page is a presentation layer on top of both. You can monitor without ever publishing a status page, and you can publish a status page that is updated manually — but automated monitoring is what makes the page trustworthy.
What monitoring actually checks
A check is a request plus an expected response. The most common types:
- HTTP/HTTPS endpoints — request a URL and verify the status code, response body, or response time. This is the default for web apps and APIs.
- DNS records — confirm a domain resolves to the expected address, catching expired records or misconfigured nameservers before users hit them.
- TCP ports — open a raw connection to a port to verify a database, mail server, or other non-HTTP service is listening.
- Custom conditions — evaluate the response against rules you define, such as "body contains
"status":"ok"" or "response time under 500 ms."
The condition matters as much as the check. A 200 response with an error message in the body is still a failure from a user's perspective, which is why body and latency conditions exist.
How automated monitoring works
The loop is simple and repeats on an interval you choose:
- Schedule — a timer triggers the check, typically every 30 seconds to 5 minutes depending on how fast you need to know.
- Request — the monitor sends the HTTP request, DNS query, or TCP connection.
- Evaluate — the response is compared against your conditions (status code, body content, latency threshold).
- Record — the result is stored, building the history that uptime percentages are calculated from.
- Alert — if the condition fails, a notification fires to email, Slack, PagerDuty, or a webhook.
Most tools add a confirmation step: a single failed check may not trigger an alert, but two or three consecutive failures will. This avoids waking someone up over a transient network blip.
Setting up basic monitoring
The exact steps depend on your tool, but the sequence is consistent:
- Pick what to monitor. Start with your most critical user-facing URL — the homepage or a health endpoint — plus one dependency like your database port or DNS record.
- Define the healthy condition. For a web app, that usually means HTTP 200 and a response body containing an expected string. For an API, check a specific field in the JSON.
- Set the interval. Shorter intervals detect problems faster but generate more traffic and noise. One minute is a reasonable starting point for most sites.
- Configure alerting. Choose a channel your team actually watches. Include the check name, the failure reason, and a link to the monitor in the alert message.
- Add a confirmation threshold. Require two or three consecutive failures before alerting to filter out false positives.
- Verify it works. Temporarily point the check at a URL you know will fail, confirm the alert arrives, then restore it.
If you are using an open-source tool like Gatus, the configuration is typically a YAML file where each endpoint gets a name, URL, interval, conditions, and alert destinations. The same structure applies whether you self-host or use a hosted service.
Common failures and how to troubleshoot them
False positives from transient errors. A single timeout during a deploy or a brief network hiccup can fire an alert. Fix: require consecutive failures before notifying.
Monitoring the wrong thing. A check that only confirms the server responds with 200 will miss a broken database connection that returns 200 with an error page. Fix: add body-content conditions, not just status codes.
Alert fatigue. If every minor latency spike pages someone, the team stops reading alerts. Fix: separate critical alerts (service down) from warnings (slow response) and route them to different channels.
Checks failing from the monitor's location but not for users. A firewall or geo-blocking rule may reject your monitor's IP. Fix: run checks from multiple regions, or whitelist the monitor's addresses.
No baseline, so you cannot tell if something is wrong. Without historical data, a 2-second response time looks fine until you realize it used to be 200 ms. Fix: let monitoring run long enough to establish normal ranges before setting latency thresholds.
The goal is not to monitor everything — it is to monitor what breaks in ways that matter, with conditions specific enough to catch real problems and thresholds loose enough to avoid crying wolf.