· via dev.to (home feed)
A leftover PGPASSWORD variable can keep a Patroni old leader from rejoining
A dev.to post shows how a PGPASSWORD variable inherited by the Patroni process makes pg_rewind fail with an authentication error, leaving the old leader stuck after a real failover.

The failure
A post on dev.to documents a Patroni failure mode that survives routine testing: a PostgreSQL cluster passes every planned switchover, but when the leader crashes for real, a replica is promoted and the old leader never rejoins. It sits in start failed in patronictl output while its PostgreSQL log repeats a timeline mismatch: the node's latest checkpoint is on timeline 2, but the cluster has forked onto timeline 3 at an earlier WAL position.
That message means the old leader wrote WAL the new leader never received, so it cannot simply follow the new timeline. Patroni's answer to this is pg_rewind, which rewinds the node's data directory to match the new primary. In the scenario the author describes, Patroni's log shows pg_rewind exiting with code 1 after the new primary rejects the connection with a password authentication failure for rewind_user. Patroni then skips the rewind and tries to start the node as a secondary, which cannot work.
PGPASSWORD beats the .pgpass file
According to the post, Patroni does not pass passwords on the command line. It writes credentials to a .pgpass file, controlled by the postgresql.pgpass setting, and points pg_rewind and pg_basebackup at it. The problem is libpq's password lookup order: a password in the connection string wins, then the PGPASSWORD environment variable, and only then the password file.
If the Patroni process inherits PGPASSWORD, typically a superuser password exported in someone's shell, pg_rewind connects as rewind_user but sends the wrong credential. Authentication fails, the rewind is skipped, and the node stays broken.
Why clean tests never catch it
The variable only harms the Patroni process that inherited it. Nodes started cleanly by systemd weeks earlier keep replicating happily. The dangerous moment is the restart after an incident: an operator logs in, exports PGPASSWORD to run a few psql commands, and starts Patroni from that same shell, or a recovery script does it for them. That process is the one that must run pg_rewind, and it carries the wrong password.
It compounds from there. Patroni points replication at the same passfile rather than embedding a password in primary_conninfo, so a node running with that environment cannot stream from the leader either, even after a manual rebuild. The author reports that in their lab every planned switchover passed while every crash test failed until the variable was found in the restart path.
The post lists four ways the variable sneaks in: manual nohup launches from a shell where PGPASSWORD was exported, deploy scripts that source an environment file with set -a, systemd units with Environment= or EnvironmentFile= entries shared with other tools, and container images that set PGPASSWORD for convenience.
Check and fix
The ten-second check is to inspect the environment of the running Patroni process on each node, via the process entry under /proc, filtering for PG and PATRONI_ prefixed variables. Any PGPASSWORD, PGUSER, PGHOST or PGSERVICE line is a problem. Unexpected PATRONI_ lines are too: Patroni reads every PATRONI_-prefixed variable as configuration, and those override patroni.yml. The author hit a related case where a PATRONI_LOG_DIR set by a script sent logs somewhere unexpected.
The fix is to run Patroni only through systemd with a clean, explicit environment: define PATH in the unit, remove PGPASSWORD from any Environment= or EnvironmentFile= lines, then daemon-reload and restart Patroni. PostgreSQL keeps running while Patroni restarts. On the stuck node, restarting Patroni is usually enough because it retries pg_rewind with the correct password. If the WAL it needs has already been recycled, rebuild the node with patronictl reinit.
Proving it is fixed
A switchover proves nothing here; only a crash does. The author's drill kills the leader's Patroni and PostgreSQL with SIGKILL while writes are in flight, waits for a new leader, restarts the old node, and confirms it rejoins as a streaming replica. After the fix, crashed leaders in their lab rejoined in about 34 seconds, consistently. The post closes by promoting the author's commercial HA kit, whose systemd units, preflight warnings and config generator are designed to prevent these traps.
Why it matters
Failover tooling is only as reliable as the environment it starts in. A single inherited environment variable can silently defeat Patroni's core recovery path, and it surfaces exactly when it hurts most: during a real incident, when a human intervenes. The lessons generalise beyond Patroni: audit process environments for PG* and PATRONI_* variables, prefer clean systemd units over manual restarts, and validate failover with hard crashes rather than graceful switchovers. The specific timings and fixes come from one lab's testing on Patroni 4.1 and PostgreSQL 16.
- #postgresql
- #patroni
- #high-availability
- #devops
- #databases