· via dev.to (home feed)
Docker Compose depends_on Doesn't Guarantee Postgres Readiness, and the Obvious Fix Isn't Enough
A dev.to write-up measured a 22% failure rate when Compose's depends_on gated app startup on Postgres, and shows why even the standard healthcheck fix leaves a socket-versus-TCP gap.

A dev.to post walks through a familiar kind of flakiness — integration tests failing with ECONNREFUSED on the first database connection, then passing on a re-run — and pins it on a specific gap in Docker Compose: depends_on does not wait for Postgres to accept connections. To rule out chance, the author tore the stack down with a fresh volume and booted it 50 times; 11 runs failed on the first connection attempt, a 22% failure rate.
What depends_on actually waits for
According to the post, the short form of depends_on waits only until the dependency container has been created and started — meaning the process inside it has launched, not that it is listening on a port or has finished initializing. Compose sees the database container running and immediately starts the dependent service, while Postgres may still be running initdb, executing init scripts, or restarting after first-time initialization.
That explains why the failure is intermittent and environment-dependent. On a developer laptop the named volume already exists, so Postgres skips initialization and is ready in well under a second. In CI, every run gets a fresh volume and pays the full init cost, which is exactly where the race shows up. The author's reproduction loop relied on docker compose down -v; without the -v flag the volume survives, initialization is skipped, and the bug appears not to exist.
The first fix, and where it stops short
The standard remedy is the long form of depends_on with condition: service_healthy, backed by a healthcheck on the database service. Compose then holds dependent services until the healthcheck passes. The post flags three details that trip people up here: $$ is needed so Compose does not interpolate ${POSTGRES_USER} from the host environment at parse time, CMD-SHELL rather than CMD so the shell can expand variables, and an explicit interval: 2s, because the default healthcheck interval is 30 seconds and can leave a stack idling for half a minute even after Postgres is ready.
After applying this, the author's failure count dropped from 11 of 50 to 2 — better, but not zero.
The remaining race: pg_isready checks a socket, the app uses TCP
The residual bug, which the author says almost nobody writes about, sits inside the official Postgres image. On a fresh volume its entrypoint runs initdb, starts a temporary server with TCP disabled — it listens only on a Unix socket — creates the configured user and database, runs anything in /docker-entrypoint-initdb.d, stops that server, and only then starts the real TCP listener.
A plain pg_isready connects over the Unix socket, so it can report success during the temporary-server phase. Docker marks the container healthy, Compose releases the app, and the app's TCP connection to port 5432 is refused because the real server has not started yet. The window is small with an empty init directory and grows with seed scripts. The fix is one flag, forcing the healthcheck onto the same path the client uses:
yaml healthcheck: test: ["CMD-SHELL", "pg_isready -h 127.0.0.1 -U $${POSTGRES_USER} -d $${POSTGRES_DB}"] interval: 2s timeout: 3s retries: 30 start_period: 10s
The temporary server never listens on TCP, so the check cannot pass until the real server is up, and start_period gives initialization grace time that does not count against retries. With that in place, the author measured 0 failures in 50 boots, and 0 again in a second batch of 50.
Ordering one-shot jobs
For migrations that must finish before the application starts, the post recommends condition: service_completed_successfully, which waits until a dependency container exits with code 0; a non-zero exit keeps the app from booting against a half-migrated schema. Compose offers three conditions in total: service_started (equivalent to the short form), service_healthy, and service_completed_successfully.
Why it matters
This failure mode is probabilistic and invisible exactly where you debug it: laptops with warm volumes almost never see it, while CI with fresh volumes sees it constantly, and it gets written off as infrastructure noise. The broader lesson from the post is that a healthcheck must exercise the same path the client uses — socket versus TCP, the right host, the right database — because a check that verifies something adjacent will pass at precisely the wrong moment. Healthchecks also gate startup order only, so the author keeps retry logic in the application for everything that happens after boot.
- #docker
- #docker-compose
- #postgresql
- #devops
- #continuous-integration