What Are Monitoring Alerts and How Do You Set Them Up?

A monitoring alert is a notification that fires when a check on your service fails or crosses a threshold you define. You set one up by picking an endpoint to watch, defining the failure condition, and choosing a channel to notify — then tuning retries and cooldowns so you get signal instead of noise. This applies to any service you can probe over HTTP, DNS, or TCP, and it's the mechanism that turns a status page from a passive dashboard into something that wakes you up.

Alert conditions vs. alert channels

These are two separate decisions, and mixing them up is the most common source of confusion.

Alert conditions answer when to fire:

Condition type Example trigger
Availability Endpoint stops responding or connection is refused
Status code Response returns 500 instead of 200
Latency Response takes longer than your threshold (e.g. 2s)
Content Response body is missing an expected string
Certificate / DNS TLS cert nearing expiry, DNS record changed or failed to resolve

Alert channels answer who finds out and how:

  • Email — universal, but easy to ignore or filter into a folder nobody reads
  • Chat (Slack, Discord, Teams) — good for team visibility, bad for 3 a.m. wake-ups
  • Webhook — posts a JSON payload to a URL you control, so you can route it into PagerDuty, an incident tool, or your own automation
  • SMS / phone — reserved for conditions that genuinely justify waking someone

A practical pattern: route availability failures to a channel that pages a human, and route latency or content warnings to a chat channel that's reviewed during working hours.

How to configure an alert

The exact menu names differ between tools, but the sequence is the same everywhere.

  1. Pick the endpoint to watch. Define the URL or host, the protocol (HTTP, DNS, TCP), and how often to check — commonly every 30s to 5 minutes. Checking every second rarely helps and can look like an attack to the target.
  2. Define the failure condition. Decide what "broken" means: non-2xx status, timeout beyond N seconds, missing expected content, or a failed DNS lookup. Be specific — "anything that isn't 200" is clearer than "unhealthy."
  3. Set the threshold and retries. A single failed check is often a blip, not an outage. Requiring 2–3 consecutive failures before firing filters out transient network hiccups at the cost of a small detection delay.
  4. Choose the channel and recipients. Wire up email, chat, or a webhook, and confirm the destination actually receives a test notification before you rely on it.
  5. Add a cooldown / re-notify interval. This controls how often you're reminded while the problem persists. Without it, a flapping service can generate hundreds of messages.
  6. Send a test alert. Most tools have a "send test" or "trigger test" action. Use it — a misconfigured webhook that silently drops payloads is worse than no alert at all.

Expected result: when the condition is met for the required number of consecutive checks, the channel receives one notification, and no further notifications until the cooldown expires or the service recovers.

Tuning to avoid false alarms and missed outages

Two failure modes pull in opposite directions, and the settings that fix one can cause the other.

  • Too noisy: lower the sensitivity by increasing consecutive-failure requirements, raising latency thresholds to realistic values, and adding cooldowns. Also check whether you're monitoring a flaky dependency rather than your own service.
  • Too quiet: an alert that never fires may be misconfigured rather than healthy. Verify the condition actually matches a real failure — for example, a check that treats any HTTP response as success will never catch a 500.

A useful habit is to deliberately break something in a staging environment and confirm the alert arrives through the full path, including any webhook or routing layer.

Why alerts fail to fire or fire too often

Symptom Likely cause
No alert during a real outage Condition too loose (e.g. accepts any response), or retries set so high the outage resolved first
Alert storm during brief blips No consecutive-failure requirement, or cooldown disabled
Notifications arrive but nobody acts Channel goes to a shared inbox or a chat channel with notifications muted
Webhook silently does nothing Endpoint returns an error or expects a different payload format; test it directly
Alert fires for a dependency you don't control Monitoring an upstream third party instead of your own endpoint

Where this fits

Alerts are the active half of monitoring; the status page is the passive half that shows current and historical state to users. A check that only updates a dashboard tells you about a problem when someone happens to look. An alert tells you when it happens. Most setups need both, and the alert configuration is what determines whether the system is actually useful at 3 a.m.

gatus.io
Create beautiful, automated status pages with advanced monitoring. Track HTTP, DNS, TCP endpoints with custom conditions and instant alerts. Built on…
journalstar.com
Read the latest Lincoln, NE news. Get the latest on events, sports, entertainment, lifestyles, and more.
trimet.org
Get arrival times, plan trips, see route maps and check service alerts, all in one place.
visualping.io
Monitor any website for changes with Visualping. Get instant alerts via email, SMS, API or Slack when a web page changes. Try it free today!