· via dev.to (home feed)
Reboot wiped /var/tmp and quietly broke Kafka Connect zstd compression
A dev.to postmortem shows how a kernel patch reboot cleared a tmpfs-backed /var/tmp, leaving a Kafka Connect worker without its pinned JVM temp directory and unable to compress records with zstd.

A healthy-looking worker that couldn't send
A postmortem published on dev.to describes how a routine kernel patch reboot left a Kafka Connect worker answering REST calls and reporting its connector as RUNNING while the underlying task failed. The patch, taking the host to kernel 6.1.188-233.386, landed on 2026-10-08 with a reboot around 16:10 UTC; by 22:18 IST that evening the team was triaging a FAILED task whose stack trace ran for several screens.
The host ran a distributed Connect worker on an AL2023-style image, with a source connector whose producer override forced zstd compression. The systemd unit deliberately pinned the JVM temp directory through JAVA_TOOL_OPTIONS=-Djava.io.tmpdir=/var/tmp/kafka — a sensible way to give temp files a known path, but with nothing wired up to recreate that path after a reboot. The package shipped no tmpfiles.d snippet, and the unit had no ExecStartPre to create and chown the directory. Because /var/tmp was tmpfs on this image, the reboot erased it, and a sibling worker in the fleet carried the same defect.
Why zstd made the failure lazy
According to the postmortem, the break only appeared when the task tried to produce records, not at startup. Kafka's zstd support runs through zstd-jni, which unpacks a platform-specific native library into java.io.tmpdir on first use. With the directory gone, File.createTempFile threw an IOException, the static initializer for the zstd classes failed, and every later compression attempt died with a NoClassDefFoundError wrapped in a ConnectException around sendRecords.
That sequence explains the misleading surface: the worker booted cleanly, REST on port 8083 stayed up, and only the data path was broken. The author's first instinct was a classpath problem, since NoClassDefFoundError usually suggests a missing jar, but the "cannot unpack" and File.createTempFile frames in the trace pointed at native code failing to write to disk. Restarting the service without creating the directory first just reproduced the crash, and restarting only the task could not recover a JVM whose initializer had already tripped. The author now trusts the task state field from the REST status endpoint plus fresh worker logs over a possibly stale trace cached in the status JSON.
Monitoring noise on port 8778
During triage, Telegraf logged connection-refused errors for Jolokia on localhost:8778 every fifteen seconds. That turned out to be unrelated: the Connect unit exposed JMX metrics through JMX Prometheus on port 7073, and 8778 existed only in the monitoring configuration. The author treated the metrics gap as a separate ticket from the data-path outage, and flagged a related pitfall — EC2 reboot timestamps and Telegraf log lines arrive in UTC while the team reasoned in IST, which distorts any timeline that isn't converted.
The fix
The immediate repair was to create the directory with kafka ownership and mode 1770, restart the service, then restart the failed task over REST. The task returned to RUNNING and downstream lag on that pipeline shard — the real impact, since the source connector could not produce until the directory existed — drained.
For durability, the postmortem recommends a tmpfiles.d entry such as d /var/tmp/kafka 1770 kafka kafka - so systemd recreates the directory on every boot, optionally backed by ExecStartPre mkdir and chown lines in the unit, with both managed through configuration management rather than left as a manual post-reboot step. Switching the producer compression override to lz4, snappy, or none would sidestep the native unpack entirely, at the cost of compression ratio or CPU; the team kept zstd and fixed the directory. Fleet-wide rollout of the tmpfiles.d change and the Telegraf cleanup were still open when the incident was closed.
Why it matters
Any JVM service that pins java.io.tmpdir to a tmpfs-backed path — or leans on a native library that unpacks itself lazily on first use — carries this failure mode, and routine patch reboots are precisely when it surfaces. Health checks that only hit REST or verify process state will keep passing while records stop flowing, and the resulting stack trace points investigators toward classpaths and brokers rather than a missing directory. The cheap insurance is a one-line tmpfiles.d snippet plus an ExecStartPre, and the cheap diagnostic is confirming the tmpdir exists before spending time elsewhere. Unit files, environment settings and filesystem layout deserve the same configuration-management rigor as the application itself.
- #kafka
- #kafka-connect
- #systemd
- #devops
- #incident-postmortem