← All docs

Heartbeat monitoring

Scheduled jobs fail quietly. Give each job a heartbeat monitor and have it call a private URL every time it finishes. If a call doesn't arrive in time, we alert you. Heartbeat monitors are on the Starter and Pro plans.

Set up a heartbeat

  1. In the app, choose Add monitor → Heartbeat.
  2. Set Expect a ping every to how often the job runs (1 minute to 7 days), and a grace period for how late it may be (up to 24 hours).
  3. Copy the monitor's ping URL. It looks like https://spotdowntime.com/api/v1/ping/YOUR_TOKEN.
  4. Call it at the end of the job, only when the job succeeded.
The token in the URL is the only key: anyone with the URL can ping the monitor. Keep it in the job's configuration, not in public code. If it leaks, delete the monitor and create a new one.

How a heartbeat decides it's down

Each ping sets a deadline: now + interval + grace period. If the next ping hasn't arrived by then, the monitor goes down, an incident opens and your alert channels are notified. The next ping brings it back up and sends a recovery alert.

The first deadline starts when you create the monitor, so a job that never runs is caught too. Pings to a paused monitor are ignored. A successful ping within 10 seconds of the previous one is ignored, so retries in your job don't pile up.

Ping URL reference

RequestEffect
GET /api/v1/ping/TOKENThe job ran. Marks the monitor up and resets the deadline.
POST /api/v1/ping/TOKENSame as GET. Any request body is ignored.
HEAD /api/v1/ping/TOKENSame as GET, for tools that can only send HEAD.
GET or POST …/TOKEN/failThe job failed. Marks the monitor down right away and alerts, without waiting for the deadline.

A successful call answers 200 with {"ok":true}. An unknown token answers 404. No login or API key is needed.

Examples

Shell script. -f makes curl fail on HTTP errors, -m 10 stops a slow network from hanging the job, and --retry 3 covers a brief blip:

bash
#!/bin/sh
set -e
/opt/backup.sh
curl -fsS -m 10 --retry 3 https://spotdowntime.com/api/v1/ping/YOUR_TOKEN > /dev/null

Crontab. && pings only if the job succeeded; || reports a failure instead:

crontab
0 2 * * * /opt/backup.sh && curl -fsS -m 10 https://spotdowntime.com/api/v1/ping/YOUR_TOKEN || curl -fsS -m 10 https://spotdowntime.com/api/v1/ping/YOUR_TOKEN/fail

GitHub Actions. Store the URL as a repository secret named HEARTBEAT_URL:

yaml
- name: Report to Spot Downtime
  if: always()
  run: |
    if [ "${{ job.status }}" = "success" ]; then
      curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL"
    else
      curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL/fail"
    fi
  env:
    HEARTBEAT_URL: ${{ secrets.HEARTBEAT_URL }}

Kubernetes CronJob.

yaml
apiVersion: batch/v1
kind: CronJob
metadata:
  name: nightly-report
spec:
  schedule: "0 3 * * *"
  jobTemplate:
    spec:
      template:
        spec:
          restartPolicy: OnFailure
          containers:
            - name: report
              image: your-image
              command: ["/bin/sh", "-c", "run-report && curl -fsS -m 10 $HEARTBEAT_URL"]
              env:
                - name: HEARTBEAT_URL
                  valueFrom:
                    secretKeyRef: { name: heartbeat, key: url }

Python.

python
import urllib.request

def heartbeat(ok: bool = True) -> None:
    url = "https://spotdowntime.com/api/v1/ping/YOUR_TOKEN" + ("" if ok else "/fail")
    try:
        urllib.request.urlopen(url, timeout=10)
    except OSError:
        pass  # never let monitoring break the job itself

Node.js.

javascript
await fetch("https://spotdowntime.com/api/v1/ping/YOUR_TOKEN", { signal: AbortSignal.timeout(10_000) }).catch(() => {});

Choosing the interval and grace period

Set the interval to the job's schedule, and the grace period to its longest normal run time plus a margin. A nightly backup that usually takes 20 minutes: interval 1 day, grace 1 hour. A queue worker that pings every minute: interval 1 minute, grace 2 minutes.

Too short a grace period gives false alarms when a job runs long; too long delays real alerts. The monitor's page shows every ping, so you can see how punctual the job really is and adjust.

Next steps

Send the alerts somewhere your team will see them: Slack, Teams, email or your own system with webhooks.

Something missing or unclear? Tell us.