How to monitor cron jobs and catch the ones that silently stop
Cron jobs fail without telling anyone. Learn the heartbeat pattern: the job checks in when it finishes, and you get alerted when it doesn't, with examples for crontab, GitHub Actions and Kubernetes.
· 5 min read · By the Spot Downtime team
Website outages are loud: customers email you, your phone buzzes, someone posts a screenshot. Cron jobs fail in silence. The nightly backup stops running, the invoice job crashes on one bad record, the cleanup script never gets deployed to the new server. Nothing breaks today. You find out weeks later, usually at the worst possible moment.
The fix is a simple pattern called a heartbeat (also called a dead man's switch). This guide explains how it works, how to set it up for the common schedulers, and how to choose settings that don't send false alarms.
Why normal uptime checks can't see cron jobs
A normal uptime monitor asks “is this URL responding?” every minute or so. A cron job has no URL to ask. It runs on its own schedule, does its work and exits. When it fails, there's often nothing to notice:
- The job never starts. The server was replaced and nobody copied the crontab, the scheduler is paused, the container image changed.
- The job starts but crashes. An expired API key, a full disk, a renamed database column. The error goes to a log file nobody reads, or to
/dev/null. - The job hangs. It waits forever on a network call and never finishes, so it never fails either.
- The job “succeeds” but does nothing. A backup script that exits 0 after writing an empty file.
Cron's built-in email on output (MAILTO) helps a little, but only if the job runs, produces output, and the server can send email that someone reads.
The heartbeat pattern
Flip the question around. Instead of asking the job whether it's alive, let the job tell you:
- You create a heartbeat monitor and get a unique URL.
- You tell the monitor how often the job runs, e.g. every day, plus a grace period for how late it may finish.
- At the very end of the job, after the real work succeeded, the job calls the URL. That's the “heartbeat”.
- If a heartbeat doesn't arrive in time, the monitor goes down and alerts you. When the next heartbeat arrives, it recovers.
Why this catches every failure mode
Setting it up
Plain crontab
Use && so the ping only happens if the job succeeded. Spot Downtime also accepts a /fail suffix, which reports a failure right away instead of waiting for the deadline:
# Nightly backup at 02:00; report success or failure
0 2 * * * /opt/backup.sh && curl -fsS -m 10 --retry 3 https://spotdowntime.com/api/v1/ping/YOUR_TOKEN \
|| curl -fsS -m 10 https://spotdowntime.com/api/v1/ping/YOUR_TOKEN/failThe curl flags matter: -f treats HTTP errors as failures, -m 10 stops a slow network from hanging your job, and --retry 3 rides out a brief network blip.
Inside a script
If the job is a script, ping from the script itself, after the work, so a crash anywhere stops the heartbeat:
import urllib.request
def main():
export_reports()
upload_to_storage()
if __name__ == "__main__":
main() # raises on failure, so the ping below never runs
try:
urllib.request.urlopen("https://spotdowntime.com/api/v1/ping/YOUR_TOKEN", timeout=10)
except OSError:
pass # never let monitoring break the jobGitHub Actions scheduled workflows
Scheduled workflows are a classic silent failure: in public repositories, GitHub disables them after 60 days without repository activity, and runs can be delayed or skipped when GitHub is busy. A heartbeat tells you when that happens:
on:
schedule:
- cron: "0 6 * * *"
jobs:
sync:
runs-on: ubuntu-latest
steps:
- run: ./sync.sh
- if: success()
run: curl -fsS -m 10 --retry 3 "${{ secrets.HEARTBEAT_URL }}"Kubernetes CronJobs
Chain the ping onto the container command, so it runs only if the work succeeds. Store the URL in a Secret rather than in the manifest:
command: ["/bin/sh", "-c", "run-report && curl -fsS -m 10 $HEARTBEAT_URL"]For more examples (Node.js, workers, failure reporting), see the heartbeat monitoring docs.
Choosing the interval and grace period
Two settings decide whether your heartbeat is useful or noisy:
- Interval: how often the job runs. For a daily job, 1 day.
- Grace period: how late it may report before you're alerted. Set it to the job's longest normal run time plus a margin.
Some sensible starting points:
- Nightly backup that usually takes 20 minutes: interval 1 day, grace 1 hour.
- Hourly report that takes a few minutes: interval 1 hour, grace 15 minutes.
- Queue worker that pings every minute while it's healthy: interval 1 minute, grace 2 to 3 minutes.
Too short a grace period sends false alarms whenever a job runs a bit long; too long delays the real alert. After a week, look at the actual ping times and tighten it.
Common mistakes
- Pinging at the start of the job. Then you only learn that the job started, not that it worked. Ping at the end.
- Pinging unconditionally.
backup.sh; curl …pings even when the backup failed. Use&&, or report/failexplicitly. - Letting monitoring break the job. Always set a timeout and ignore errors from the ping itself.
- One heartbeat for many jobs. Give every important job its own monitor, so the alert tells you exactly which one stopped.
- Committing the URL to a public repo. The token in the URL is the key. Keep it in a secret or environment variable.
Which jobs deserve a heartbeat?
Start with the ones where silence is expensive:
- Backups and database dumps
- Billing, invoicing and payouts
- Data syncs and imports from partners
- Certificate renewal and other housekeeping that keeps the site up
- Email digests and notifications customers expect
- Long-running queue workers and consumers
Heartbeat monitors are part of Spot Downtime's Starter and Pro plans, alongside the website, API and DNS checks. See how cron job monitoring works, or create a free account to set up your first monitors in a couple of minutes.
Keep reading
- Uptime · SLAsWhat 99.9% uptime really means (with a downtime table)99.9% sounds close to perfect, but it allows 43 minutes of downtime a month. Here's what each common uptime target allows, why your real uptime is lower than any one provider's, and how to pick an SLA you can keep.October 2, 2026 · 5 min read
- TroubleshootingWebsite down? A 10-minute checklist to find the causeA calm, step-by-step way to find out why a site is down: is it really down, DNS, the certificate, the server, or the app. Each step takes a minute and says what to do next.October 2, 2026 · 5 min read