502 vs 503 vs 504: what each error means and how to fix it

Three gateway errors, three different problems: a crashed app, no capacity, or a slow backend. How to tell them apart, where to look first, and what Cloudflare's 52x codes mean.

· 4 min read · By the Spot Downtime team

502, 503 and 504 look alike: three numbers, one meaning for your visitors (“the site is broken”). For you, they point in different directions. Each one tells you which part of the chain failed, and knowing which saves a lot of guessing in the middle of an outage.

First: where these errors come from

Most sites aren't one server. A request passes through a chain: a CDN or load balancer, then a reverse proxy like nginx, then your application, which may call a database or other APIs. 502, 503 and 504 are usually produced by something in front of your application, describing what happened when it tried to reach the next link in the chain.

CodeWhat the proxy is sayingIn one word
502 Bad GatewayI reached the server behind me, but its answer was broken or the connection died.Crashed
503 Service UnavailableThe service can't handle requests right now: overloaded, out of capacity, or in maintenance.Unavailable
504 Gateway TimeoutI asked the server behind me and it didn't answer in time.Slow

502 Bad Gateway: the app crashed or isn't listening

The proxy got through to your application, or tried to, and got nonsense or nothing back. Common causes:

  • The application process crashed, or is restarting in a loop.
  • It ran out of memory and the system killed it.
  • The proxy points at the wrong port or socket after a deploy, so it connects to nothing (often logged as connect() failed (111: Connection refused) in nginx).
  • The app sends a response the proxy can't parse, such as oversized headers.

Where to look

bash
sudo tail -n 50 /var/log/nginx/error.log      # what the proxy saw
systemctl status your-app                     # is the app running?
journalctl -u your-app --since "15 min ago"   # why it stopped
dmesg -T | grep -i "out of memory"            # killed for memory?

If a deploy just went out, a 502 is very often the new version failing to start. Roll back first.

503 Service Unavailable: no capacity, or on purpose

A 503 means “try again later”. Sometimes that's intentional, like a maintenance page, and sometimes the system has no capacity left:

  • No healthy servers. Load balancers return 503 when every server behind them fails its health check.
  • Overload. All workers are busy, the queue is full, or a rate limit kicked in.
  • Maintenance mode. Your platform or app is deliberately refusing requests.
  • A dependency is down and your app reports itself unavailable, which is the right thing to do.

Where to look

Check the load balancer's view of your servers' health first, then CPU, memory and worker counts. A 503 that started with a traffic spike points to capacity; one that started with a deploy points to health checks the new version fails.

Doing maintenance? Send a 503 on purpose

For planned downtime, a 503 with a Retry-After header is the correct answer: search engines know the outage is temporary and don't drop your pages. A 200 maintenance page tells them your real content is gone.

504 Gateway Timeout: the app is too slow

The proxy waited for your application and gave up. The app may be fine; it just didn't finish in time. nginx waits 60 seconds by default (proxy_read_timeout); load balancers and CDNs have their own limits.

  • A slow database query, often a missing index on a table that grew.
  • A call to an external API that hangs, with no timeout set in your code.
  • Locks: one long transaction blocking everything behind it.
  • A request doing too much work, like a large export or report, inline instead of in a background job.

Where to look

Find the slow requests in your application logs or APM, then the slow queries in the database's slow query log. Raising the proxy timeout hides the symptom; fixing the slow part, or moving it to a background job, fixes the problem.

Behind Cloudflare? You'll see 52x instead

Cloudflare uses its own codes for the same situations, which tell you the problem is between Cloudflare and your server:

CodeMeaning
521Your server refused the connection: the web server is down, or a firewall blocks Cloudflare.
522The connection to your server timed out.
523Cloudflare can't reach your server: a DNS or routing problem with the origin.
524Connected, but your server didn't respond within Cloudflare's time limit, like a 504.
525The SSL handshake with your server failed.
526Your server's certificate isn't valid for Cloudflare's strict SSL mode.

Intermittent errors are still errors

The hardest version of all three is the one that happens for a minute and goes away: a worker that crashes under load, a query that's slow only at peak times. By the time someone looks, everything works.

That's where monitoring earns its keep. An uptime monitor records every failed check with the status code and error, so instead of “the site was flaky this morning” you get “502s for 3 minutes at 09:14, right after the deploy”. Spot Downtime keeps that history per monitor, alongside response times, so slow-before-it-breaks patterns show up too.

Website down checkerCheck any URL from outside right now and see the exact status code and error it returns.

Working through an outage? The website down checklist walks through the whole diagnosis, step by step.

Keep reading