· via dev.to (home feed)
A source-verified map of where NGINX 502 and 504 errors actually originate
A dev.to write-up models NGINX 1.30 as a state machine checked against the source code, pinpointing exactly where 502 and 504 responses are generated and which defaults decide between them.

A state-machine map of NGINX request handling
A write-up on dev.to tackles two of the most familiar production failures from an unusual angle: instead of listing fixes, the author models open-source NGINX stable 1.30 as a hierarchical state machine and then verifies every transition against the official nginx.org documentation, the CHANGES file and the source code itself. The payoff is a precise account of where 502 Bad Gateway and 504 Gateway Timeout responses are actually produced, and which configuration defaults determine whether you see one or the other.
The 502 versus 504 split
According to the post, the divergence happens in a single function, ngx_http_upstream_next(). Timeouts map to 504: NGINX was waiting on the upstream and the wait exceeded the configured limits. Connection errors, invalid response headers and the "no live upstreams" condition map to 502: NGINX could not obtain a usable response at all.
That distinction is the debugging shortcut. A 504 sends you looking at backend latency and timeout tuning; a 502 sends you checking whether the upstream is reachable, healthy and returning valid HTTP.
Eleven phases and a bounded rewrite loop
Request processing, as the author counts it, runs through eleven phases. When a rewrite changes the URI, control loops back to the FIND_CONFIG phase, and NGINX allows that cycle at most ten iterations before returning a 500 error. Configurations with long rewrite chains therefore hit a hard ceiling rather than looping forever.
New defaults in stable 1.30
One change is easy to miss when reading older material. Since 1.29.7, and therefore in stable 1.30, proxy keepalive and HTTP/1.1 are the default for upstream connections, and the Connection header is no longer sent. Debugging intuition built on the older model, where each proxied request got a fresh connection, may no longer match what actually happens between NGINX and the backend.
Retries and upstream health
Whether NGINX tries another server after an upstream failure depends on proxy_next_upstream, which the post says defaults to "error timeout". There is a carve-out for non-idempotent methods: requests using POST, LOCK or PATCH that have already been transmitted are not retried, so a single failed attempt can be the whole story for those.
Upstream health is passive by default, with max_fails set to 1 and fail_timeout to 10 seconds. The active health_check directive the post points to belongs to the commercial NGINX Plus product rather than the open-source build.
Startup and reload behaviour
Two operational details round out the picture. On startup, a failed bind() is retried five times, 500 milliseconds apart, before NGINX reports that it still could not bind. During a reload, the master process validates and applies the new configuration first: if that fails, the previous configuration stays in place and existing workers keep serving; if it succeeds, new workers start while old workers finish the requests they are already handling. A well-behaved reload is thus designed not to drop in-flight work.
Why it matters
502 and 504 are symptoms, not causes, and they point at different causes. Knowing that timeouts become 504 while connection failures and malformed responses become 502 turns a vague incident report into a directed investigation. The defaults matter just as much: retry rules that skip already-sent POSTs, passive health checks that can sideline a server after one failure within ten seconds, and newly default keepalive connections all change which error appears and when. For teams running NGINX in front of flaky backends, this kind of source-verified map is worth reading before the next pager alert rather than during it.
- #nginx
- #debugging
- #web-server
- #http
- #devops