· via Hacker News – Front Page (native)
Swap pushed Go GC metadata to disk, causing 40ms stop-the-world pauses
A developer's swap experiment turned typical 51-microsecond Go GC pauses into 40ms stop-the-world stalls, with page faults on collector metadata to blame

How swap broke a latency budget
A developer who enabled swap in production to absorb memory spikes has documented how the decision produced 40ms stop-the-world pauses in a Go service — roughly 800 times the typical pause. In a technical write-up on frn.sh that reached Hacker News's front page, the author traces the stalls to page faults on the garbage collector's own bookkeeping.
The test environment was a cgroup holding two processes: a Go program that reads data with io.ReadAll and decodes it with proto.Unmarshal, building a blob plus a graph structure the allocator marks as containing pointers, and a mostly idle HTTP server.
The author expected the interplay to be benign. Because the kernel accounts for and evicts pages per cgroup rather than per process, reclaim pressure should spread across both processes' memory, making it unlikely the collector's own pages would be singled out. Testing proved that reasoning wrong.
The metadata outside the heap
The finding at the center of the post: Go's collector reads its metadata during stop-the-world pauses, and that metadata lives outside the heap in pages the runtime allocates once, reuses and never frees. Nothing stops the kernel from evicting them.
Since reclaim targets the least recently used pages, and the GC touches its bookkeeping only in bursts between cycles, those pages age out. The next time the collector stops the world and reads them, it triggers major page faults: the kernel walks page tables, calls do_swap_page, finds and charges a new frame, submits I/O and waits for the disk — while every Go scheduler slot sits halted.
The numbers
On a Hetzner machine running kernel 6.8 with MGLRU enabled, the median pause measured about 51 microseconds. With metadata paged out to the local NVMe, the worst pause reached 40ms.
A BPF script that counts page faults while the world is stopped explained the gap: the worst pause, at 39,902 microseconds, spent 39,013 microseconds inside 228 page faults, with stack traces landing in GC bookkeeping such as spanSet.reset, finishsweep_m, nextMarkBitArenaEpoch and gcStart. Go stops the world at sweep termination and mark termination; the 30-minute run recorded 312 such pauses, with the worst-case stalls appearing two to three times per memory spike.
A second, localized cost
Separately, building a single 511 KiB message — normally a 3-5ms operation — took 105ms with the data on NVMe and 903ms on Hetzner's network-backed volume. Per message that exceeds the metadata pause cost, though only the allocating goroutine pays it rather than the entire process. The author has not yet confirmed where that time goes.
Asked whether Go 1.26's Green Tea garbage collector changed how metadata is read, the author measured it and found the impact negligible. Plots, a mock allocator, BPF scripts and other reproduction materials are available in a companion repository.
Why it matters
A 40ms stop-the-world pause halts every goroutine, including ones waiting on I/O that completes during the stall — a visible event for latency-sensitive services. The post is also a counterweight to the blanket reassurance that swap is harmless on modern kernels: the author still agrees with kernel developer Chris Down that swap is not evil, but a tracing collector that must atomically read evictable metadata outside the heap is a specific failure mode. For allocation-heavy Go workloads — the author's own production case triggers collection frequently — the lesson is to treat swap as a variable that can quietly convert microseconds into milliseconds.
- #go
- #garbage-collection
- #linux
- #swap
- #performance