deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

What Docker HEALTHCHECK actually tests: one exit code and its limits

A dev.to walkthrough explains that Docker HEALTHCHECK turns on a single exit code, what its three health states mean, and where liveness checks stop telling you anything useful.

What Docker HEALTHCHECK actually tests: one exit code and its limits

A post on dev.to by jtorchia, originally published on juanchi.dev, breaks down what a Docker HEALTHCHECK actually does: it answers exactly one question — whether the command you configured exited with code 0 at the moment it ran. Everything else engineers assume about "health" has to be built into that command by hand.

What the command actually runs

Docker's own documentation describes HEALTHCHECK as a way to tell Docker how to test that a container is "still working", and its canonical example is a curl against the root path with a three-second timeout. As the dev.to article points out, that check proves one thing only: the web server answered something at that path within the timeout. It says nothing about whether the database is up, whether message queue workers are alive, or whether the authentication service the application depends on is responding. If the endpoint returns any response at all — including a misconfigured error page served with a 200 status — the check passes.

Docker never inspects the response body or validates application logic. The exit code alone drives the outcome: 0 marks the container healthy, 1 marks it unhealthy, and 2 is reserved. A /health endpoint that returns a hardcoded JSON payload with a 200 status will keep a container marked healthy forever, even when the real application logic behind it is broken.

Liveness versus real health

The distinction the article draws is between shallow liveness and real health. A liveness check confirms the process is responsive: it isn't hung, it hasn't locked into an infinite loop, the port is listening. Docker's documentation frames the feature exactly this way, citing the case of a web server stuck in a loop that can no longer handle new connections even though the server process is still running.

That is the full extent of the coverage. If the process answers but has lost its database connection, a check that only curls the root path will never find out — the server keeps returning 200 on that path while every real request to the application fails. A healthy status in docker ps can coexist with an application that cannot complete a single useful operation.

start_period and retries

Two HEALTHCHECK options control how patient Docker is before flagging a container unhealthy: --start-period and --retries. According to the documentation the article cites, failures during the start period do not count toward the retry limit, giving slow-booting containers room to initialise. But there is a subtlety: once the check passes even once during the start period, the container counts as started, and every consecutive failure from that point counts for real.

The consequence is that a short start_period paired with a slow-booting application can reclassify "still starting up" as a genuine failure, producing a false unhealthy status. The right values depend entirely on the application's actual bootstrap time — there is no universal setting that works for every image.

What even a good check cannot measure

The article lists limits that no HEALTHCHECK command escapes on its own:

  • External dependencies the command never touches. If it does not query the database, the healthcheck knows nothing about it.
  • Partial degradation. A slow response inside the timeout passes exactly like an instant one; the exit code has no gradients.
  • Business state. A 200 response with corrupted or empty content passes if the check only validates the HTTP status.
  • Anything outside the container. HEALTHCHECK runs inside it, so it cannot see the host, the wider network, or neighbouring containers unless the command queries them directly.

Each of those gaps is closed by adding logic to the check itself, not by changing how Docker interprets the result. If the check needs to reflect a service's real condition, the work goes into writing an endpoint or script that exercises the dependencies that matter — not into tweaking intervals or retry counts.

Why it matters

Health status feeds real operational decisions: orchestrators, load balancers and anyone reading docker ps treat "healthy" as a signal worth trusting. The core warning of the dev.to post is that this signal is only as deep as the command behind it. A green status can mean little more than "the process answered on a port". Teams shipping containers should treat HEALTHCHECK as a mechanism rather than a guarantee, and invest in check scripts that actually exercise database connections, queues and critical external services — or accept that the status column is measuring something far narrower than the word suggests.

  • #docker
  • #containers
  • #devops
  • #health-checks

Related posts