Website Review
What is Gatus?
Gatus is an open-source tool for building automated status pages. It monitors endpoints — HTTP, DNS and TCP — applies custom conditions to decide whether each one is healthy, and turns the results into a status page with alerting when something fails. It is aimed at developers and operations teams who want uptime and service-health reporting they control themselves, rather than a hosted status-page service.
What it actually does
- Checks services on a schedule. Each configured endpoint is probed repeatedly, so you get a continuous picture of availability rather than a one-off test.
- Evaluates conditions, not just reachability. A check can pass or fail based on the response itself — status codes, body content, response time — so "the server answered" and "the service works" are treated as different things.
- Publishes a status page. Results are rendered as a page your users or colleagues can read, showing which services are up, degraded or down.
- Sends alerts. When a condition breaks, Gatus notifies you through your chosen channels, which is what makes it operational rather than just decorative.
Who it suits
A small engineering team running a handful of APIs and internal services is the clearest fit: you get a public or internal status page without paying per-seat for a hosted product, and you can extend checks in code. Teams that want zero maintenance, or that need polished stakeholder-facing communication features, may find a hosted alternative less work.
Trade-offs to weigh
Self-hosting means you own deployment, upgrades, storage of check history and the alerting integrations. That is the price of flexibility: you decide what is monitored and how failures are judged, but you also carry the operational burden. A useful decision rule is whether your team already runs its own infrastructure comfortably — if yes, Gatus fits naturally; if no, the setup cost may outweigh the control.
Next step
Write down three to five checks that represent real user-facing behaviour (for example, a login endpoint returning the expected response, not just an HTTP 200), then configure those first. Starting from meaningful conditions rather than every host you own keeps the status page honest and the alerts worth reading.
How does Gatus compare to other status page tools like Statuspage or Cachet?
Gatus is aimed at developers who want monitoring and the public status page to come from the same configuration, rather than at teams that want a fully hosted, point-and-click service. Statuspage is a hosted, commercial product built around communicating incidents to customers; Cachet is an open-source, self-hosted status page whose focus is the page itself. Gatus's distinguishing idea is that the checks defined for monitoring are what drive the status page, so a service's health and its public display stay in sync without duplicating work.
H3. Where each fits
| Tool | Model | Monitoring | Status page | Best for |
|---|---|---|---|---|
| Gatus | Self-hosted, config-driven | Built in (HTTP, DNS, TCP, custom conditions, alerts) | Generated from the same checks | Developers who want one config for checks and page |
| Statuspage | Hosted, commercial | Not the core purpose | Polished incident communication | Customer-facing comms and support workflows |
| Cachet | Self-hosted, open source | Limited; page-centric | The main feature | Teams that mainly need a self-hosted page |
H3. Practical trade-offs
- Configuration vs. UI. Gatus suits people comfortable editing config files and running a service. Statuspage suits teams that want non-engineers to post updates during an incident. Cachet sits closer to Gatus in self-hosting effort but is less about automated checks.
- Automation. With Gatus, a failing endpoint can flip the page and trigger alerts automatically, which reduces the gap between "the monitor noticed" and "the page shows it." With a page-first tool, that link is usually manual.
- Incident communication. If your main problem is telling customers what is happening, when, and why, a dedicated communication tool handles that narrative better than an automated page.
- Control and cost model. Self-hosted options keep data on your infrastructure and avoid per-seat or per-page fees, but you own uptime, upgrades and alert routing. Hosted options shift that burden to the vendor.
A useful decision criterion: if your team already writes monitoring checks as code and you want the page to be a byproduct, Gatus is the natural fit. If your bottleneck is incident messaging to non-technical customers, choose a communication-first tool. If you simply want a self-hosted page with minimal automation, Cachet is the lighter option.
For a concrete next step, list the endpoints you already check, then ask whether each should appear publicly. That list tells you whether you need Gatus-style automated status or a page where humans write updates. You can see the project at Gatus.
How do I set up Gatus for monitoring my own services?
Gatus is a self-hosted, configuration-driven status page. You define your services in a YAML config file, run the Gatus binary or container, and it continuously probes those endpoints and renders a public status page with the results. There is no point-and-click setup wizard — the config file is the setup.
The basic workflow
- Install Gatus as a binary or Docker container on a host that can reach your services.
- Write a config file listing your endpoints. Each entry needs a name, a URL, and conditions that define "healthy."
- Start Gatus and visit its web UI to see the status page.
- Add alerting (email, Slack, PagerDuty, etc.) so you hear about failures without watching the page.
A minimal config looks roughly like this in structure:
endpoints:
- name: My API
url: "https://api.example.com/health"
interval: 1m
conditions:
- "[STATUS] == 200"
- "[BODY].status == UP"
- "[RESPONSE_TIME] < 500"
What Gatus actually monitors
The product supports more than HTTP checks, which matters if your stack isn't purely web-facing:
- HTTP/HTTPS — status codes, response body contents (JSON path checks), response time thresholds, and header/body assertions.
- TCP — whether a port accepts a connection, useful for databases, message brokers, or SSH.
- DNS — whether a hostname resolves, and optionally to what.
- ICMP — basic reachability of a host.
That mix makes Gatus a reasonable fit for a small team that wants one dashboard covering a web app, its database port, and its DNS record, rather than three separate tools.
A concrete reader scenario
Say you run a SaaS side project on a single VPS: a web app, a Postgres database, and an API used by a mobile client. You'd add three endpoints — an HTTP check on the app's health route, a TCP check on the database port, and an HTTP check on the API with a body assertion. Set interval to something like 60 seconds, and use conditions to distinguish "the process is up" from "the process is up and returning correct data." That second distinction is where condition-based checks earn their keep; a 200 response with an error payload still looks healthy to a naive uptime monitor.
Trade-offs to weigh
| Consideration | What it means for you |
|---|---|
| Self-hosted | You control data and cost, but you also own uptime, backups, and upgrades. If your monitoring host dies, you may not be alerted. |
| Config-file driven | Excellent for version control and reproducibility; less friendly if non-technical teammates need to add checks. |
| Condition-based checks | More precise than plain uptime, but you must know what a healthy response body looks like. |
| Status page included | You get a public-facing page without a separate service, but you should think about what you expose publicly. |
Practical next step
Start with one endpoint you already know the healthy response for — your own website's homepage or a /health route — and get the status page rendering before adding alerting. Once that works, add a second check that fails deliberately (point it at a URL returning 500) to confirm your alert channel actually fires. Verifying the alert path is the step most people skip and later regret.
If you want to compare approaches before committing, it's worth looking at hosted alternatives that trade control for zero maintenance, such as UptimeRobot or Better Stack, and at GitHub for Gatus's own repository if you want to read the config reference and community examples directly.
What types of endpoints can Gatus monitor and under what conditions?
Gatus monitors the endpoint types you would expect from a developer-focused status page tool: HTTP/HTTPS services, DNS records, and raw TCP connections. Each check is defined in configuration with a target and a set of conditions that decide whether the result counts as healthy.
Endpoint types and typical checks
- HTTP/HTTPS — request a URL and assert on the response. Conditions can cover status codes, response body contents, headers, and response time thresholds. Useful for APIs, web apps, and health endpoints.
- DNS — query a hostname and validate the answer, such as whether a record resolves to an expected value or resolves at all. Useful for catching resolver or zone problems independently of your web tier.
- TCP — open a connection to a host and port to confirm the service accepts connections. Useful for databases, message brokers, and other non-HTTP services.
How conditions work
Conditions are the core of Gatus: instead of only "is it up?", you define assertions like "status is 200", "body contains a known string", or "response time is under a limit". A check passes only when its conditions hold, so a service returning an error page with a 200 status can still be marked as failing if you assert on body content. This makes the status page reflect real availability rather than just reachability.
Choosing what to monitor
For a public-facing API, combine an HTTP check on a health route with a DNS check on the same hostname — that separates "DNS is broken" from "the app is broken". For internal infrastructure, TCP checks on ports plus HTTP checks on admin endpoints give fast, low-cost coverage. Keep conditions specific enough to catch partial failures but not so strict that normal variation triggers alerts.
A practical next step is to start with one HTTP check per critical service, add a TCP check for each non-HTTP dependency, and tighten conditions only after you see what normal responses look like.
How can I integrate Gatus with my existing alerting systems like Slack or PagerDuty?
Gatus supports alerting providers as part of its configuration, so integration with Slack or PagerDuty is done by declaring the provider and its credentials in the same config file that defines your endpoints and conditions. The status page and the alerting layer come from one tool, which is the main trade-off: you get unified monitoring without a separate alert router, but you configure alerts in Gatus rather than in a central incident platform.
A typical setup
- Define your endpoints (HTTP, DNS, TCP, ICMP) with conditions and intervals.
- Add an alerting block for each provider you use, with the webhook URL or API key.
- Attach alerts to endpoints, or set them as defaults so every new check inherits them.
- Use failure and success thresholds so a single blip doesn't page anyone; require a condition to fail several times before alerting.
- Send a recovery notification when the condition passes again, so on-call knows the incident is over.
Slack vs PagerDuty
| Slack | PagerDuty | |
|---|---|---|
| Best for | Team visibility, quick triage | On-call rotation, escalation, paging |
| Alert style | Channel message, easy to discuss | Incident with acknowledgement and escalation |
| Watch out for | Noise if thresholds are too tight | Duplicate incidents if both tools alert on the same check |
A practical pattern is to send everything to Slack for awareness and route only customer-facing or high-severity checks to PagerDuty. That keeps the rotation meaningful and prevents alert fatigue.
Things that usually go wrong
- Credentials committed to a public repository; use environment variables or a secret store.
- Alerts firing during deploys or maintenance windows; suppress or raise thresholds during known changes.
- Testing in production: trigger a deliberate failure on a staging endpoint first to confirm the payload arrives and reads clearly.
If you already run a status page elsewhere, Atlassian Statuspage and Instatus are common alternatives, but they don't replace Gatus's own monitoring and alerting config.
Next step: open a test channel or a PagerDuty service, point one low-risk endpoint at it, and confirm both the firing and recovery messages before rolling alerts out across your checks.
Is Gatus free to use and what are the licensing terms?
Gatus is free to use. It is an open-source project, so the core software is available at no cost and its licensing is governed by the project's open-source license rather than a paid subscription or per-seat plan. You can self-host it and run it on your own infrastructure without paying a license fee.
What "free" means here
- No purchase or subscription is required to obtain and run the software.
- The cost you should plan for is operational: a server or container to run it on, and the time to configure and maintain it.
- If you want managed hosting, that would come from a separate provider, not from the open-source project itself.
Practical next step
If license terms matter for your organization, check the LICENSE file in the project's repository before deploying. For most self-hosted monitoring and status-page use, the open-source license is the relevant document; if your legal team needs to confirm redistribution or commercial-use rights, that file is the authoritative source.
For a concrete scenario: a small team can run Gatus on a cheap VPS, point it at their HTTP and TCP endpoints, and publish a status page without any software fee. The trade-off is that you own uptime, upgrades and alert routing yourself, which is exactly what you would otherwise pay a hosted status-page vendor to handle.
User reviews (0)