How to monitor cron jobs and catch the ones that silently stop

Cron jobs fail without telling anyone. Learn the heartbeat pattern: the job checks in when it finishes, and you get alerted when it doesn't, with examples for crontab, GitHub Actions and Kubernetes.

· 5 min read · By the Spot Downtime team

Website outages are loud: customers email you, your phone buzzes, someone posts a screenshot. Cron jobs fail in silence. The nightly backup stops running, the invoice job crashes on one bad record, the cleanup script never gets deployed to the new server. Nothing breaks today. You find out weeks later, usually at the worst possible moment.

The fix is a simple pattern called a heartbeat (also called a dead man's switch). This guide explains how it works, how to set it up for the common schedulers, and how to choose settings that don't send false alarms.

Why normal uptime checks can't see cron jobs

A normal uptime monitor asks “is this URL responding?” every minute or so. A cron job has no URL to ask. It runs on its own schedule, does its work and exits. When it fails, there's often nothing to notice:

  • The job never starts. The server was replaced and nobody copied the crontab, the scheduler is paused, the container image changed.
  • The job starts but crashes. An expired API key, a full disk, a renamed database column. The error goes to a log file nobody reads, or to /dev/null.
  • The job hangs. It waits forever on a network call and never finishes, so it never fails either.
  • The job “succeeds” but does nothing. A backup script that exits 0 after writing an empty file.

Cron's built-in email on output (MAILTO) helps a little, but only if the job runs, produces output, and the server can send email that someone reads.

The heartbeat pattern

Flip the question around. Instead of asking the job whether it's alive, let the job tell you:

  1. You create a heartbeat monitor and get a unique URL.
  2. You tell the monitor how often the job runs, e.g. every day, plus a grace period for how late it may finish.
  3. At the very end of the job, after the real work succeeded, the job calls the URL. That's the “heartbeat”.
  4. If a heartbeat doesn't arrive in time, the monitor goes down and alerts you. When the next heartbeat arrives, it recovers.

Why this catches every failure mode

A job that never starts, crashes, hangs or is deleted has one thing in common: it doesn't reach the last line. Silence is the signal.

Setting it up

Plain crontab

Use && so the ping only happens if the job succeeded. Spot Downtime also accepts a /fail suffix, which reports a failure right away instead of waiting for the deadline:

crontab
# Nightly backup at 02:00; report success or failure
0 2 * * * /opt/backup.sh && curl -fsS -m 10 --retry 3 https://spotdowntime.com/api/v1/ping/YOUR_TOKEN \
  || curl -fsS -m 10 https://spotdowntime.com/api/v1/ping/YOUR_TOKEN/fail

The curl flags matter: -f treats HTTP errors as failures, -m 10 stops a slow network from hanging your job, and --retry 3 rides out a brief network blip.

Cron expression explainerNot sure what */15 9-17 * * 1-5 means? Paste it in to see it in plain English, with the next run times.

Inside a script

If the job is a script, ping from the script itself, after the work, so a crash anywhere stops the heartbeat:

python
import urllib.request

def main():
    export_reports()
    upload_to_storage()

if __name__ == "__main__":
    main()  # raises on failure, so the ping below never runs
    try:
        urllib.request.urlopen("https://spotdowntime.com/api/v1/ping/YOUR_TOKEN", timeout=10)
    except OSError:
        pass  # never let monitoring break the job

GitHub Actions scheduled workflows

Scheduled workflows are a classic silent failure: in public repositories, GitHub disables them after 60 days without repository activity, and runs can be delayed or skipped when GitHub is busy. A heartbeat tells you when that happens:

yaml
on:
  schedule:
    - cron: "0 6 * * *"
jobs:
  sync:
    runs-on: ubuntu-latest
    steps:
      - run: ./sync.sh
      - if: success()
        run: curl -fsS -m 10 --retry 3 "${{ secrets.HEARTBEAT_URL }}"

Kubernetes CronJobs

Chain the ping onto the container command, so it runs only if the work succeeds. Store the URL in a Secret rather than in the manifest:

yaml
command: ["/bin/sh", "-c", "run-report && curl -fsS -m 10 $HEARTBEAT_URL"]

For more examples (Node.js, workers, failure reporting), see the heartbeat monitoring docs.

Choosing the interval and grace period

Two settings decide whether your heartbeat is useful or noisy:

  • Interval: how often the job runs. For a daily job, 1 day.
  • Grace period: how late it may report before you're alerted. Set it to the job's longest normal run time plus a margin.

Some sensible starting points:

  • Nightly backup that usually takes 20 minutes: interval 1 day, grace 1 hour.
  • Hourly report that takes a few minutes: interval 1 hour, grace 15 minutes.
  • Queue worker that pings every minute while it's healthy: interval 1 minute, grace 2 to 3 minutes.

Too short a grace period sends false alarms whenever a job runs a bit long; too long delays the real alert. After a week, look at the actual ping times and tighten it.

Common mistakes

  • Pinging at the start of the job. Then you only learn that the job started, not that it worked. Ping at the end.
  • Pinging unconditionally. backup.sh; curl … pings even when the backup failed. Use &&, or report /fail explicitly.
  • Letting monitoring break the job. Always set a timeout and ignore errors from the ping itself.
  • One heartbeat for many jobs. Give every important job its own monitor, so the alert tells you exactly which one stopped.
  • Committing the URL to a public repo. The token in the URL is the key. Keep it in a secret or environment variable.

Which jobs deserve a heartbeat?

Start with the ones where silence is expensive:

  • Backups and database dumps
  • Billing, invoicing and payouts
  • Data syncs and imports from partners
  • Certificate renewal and other housekeeping that keeps the site up
  • Email digests and notifications customers expect
  • Long-running queue workers and consumers

Heartbeat monitors are part of Spot Downtime's Starter and Pro plans, alongside the website, API and DNS checks. See how cron job monitoring works, or create a free account to set up your first monitors in a couple of minutes.

Keep reading