Heartbeat monitoring
Set up a heartbeat
- In the app, choose Add monitor → Heartbeat.
- Set Expect a ping every to how often the job runs (1 minute to 7 days), and a grace period for how late it may be (up to 24 hours).
- Copy the monitor's ping URL. It looks like
https://spotdowntime.com/api/v1/ping/YOUR_TOKEN. - Call it at the end of the job, only when the job succeeded.
How a heartbeat decides it's down
Each ping sets a deadline: now + interval + grace period. If the next ping hasn't arrived by then, the monitor goes down, an incident opens and your alert channels are notified. The next ping brings it back up and sends a recovery alert.
The first deadline starts when you create the monitor, so a job that never runs is caught too. Pings to a paused monitor are ignored. A successful ping within 10 seconds of the previous one is ignored, so retries in your job don't pile up.
Ping URL reference
| Request | Effect |
|---|---|
GET /api/v1/ping/TOKEN | The job ran. Marks the monitor up and resets the deadline. |
POST /api/v1/ping/TOKEN | Same as GET. Any request body is ignored. |
HEAD /api/v1/ping/TOKEN | Same as GET, for tools that can only send HEAD. |
GET or POST …/TOKEN/fail | The job failed. Marks the monitor down right away and alerts, without waiting for the deadline. |
A successful call answers 200 with {"ok":true}. An unknown token answers 404. No login or API key is needed.
Examples
Shell script. -f makes curl fail on HTTP errors, -m 10 stops a slow network from hanging the job, and --retry 3 covers a brief blip:
#!/bin/sh
set -e
/opt/backup.sh
curl -fsS -m 10 --retry 3 https://spotdowntime.com/api/v1/ping/YOUR_TOKEN > /dev/nullCrontab. && pings only if the job succeeded; || reports a failure instead:
0 2 * * * /opt/backup.sh && curl -fsS -m 10 https://spotdowntime.com/api/v1/ping/YOUR_TOKEN || curl -fsS -m 10 https://spotdowntime.com/api/v1/ping/YOUR_TOKEN/failGitHub Actions. Store the URL as a repository secret named HEARTBEAT_URL:
- name: Report to Spot Downtime
if: always()
run: |
if [ "${{ job.status }}" = "success" ]; then
curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL"
else
curl -fsS -m 10 --retry 3 "$HEARTBEAT_URL/fail"
fi
env:
HEARTBEAT_URL: ${{ secrets.HEARTBEAT_URL }}Kubernetes CronJob.
apiVersion: batch/v1
kind: CronJob
metadata:
name: nightly-report
spec:
schedule: "0 3 * * *"
jobTemplate:
spec:
template:
spec:
restartPolicy: OnFailure
containers:
- name: report
image: your-image
command: ["/bin/sh", "-c", "run-report && curl -fsS -m 10 $HEARTBEAT_URL"]
env:
- name: HEARTBEAT_URL
valueFrom:
secretKeyRef: { name: heartbeat, key: url }Python.
import urllib.request
def heartbeat(ok: bool = True) -> None:
url = "https://spotdowntime.com/api/v1/ping/YOUR_TOKEN" + ("" if ok else "/fail")
try:
urllib.request.urlopen(url, timeout=10)
except OSError:
pass # never let monitoring break the job itselfNode.js.
await fetch("https://spotdowntime.com/api/v1/ping/YOUR_TOKEN", { signal: AbortSignal.timeout(10_000) }).catch(() => {});Choosing the interval and grace period
Set the interval to the job's schedule, and the grace period to its longest normal run time plus a margin. A nightly backup that usually takes 20 minutes: interval 1 day, grace 1 hour. A queue worker that pings every minute: interval 1 minute, grace 2 minutes.
Too short a grace period gives false alarms when a job runs long; too long delays real alerts. The monitor's page shows every ping, so you can see how punctual the job really is and adjust.
Next steps
Send the alerts somewhere your team will see them: Slack, Teams, email or your own system with webhooks.
Something missing or unclear? Tell us.