deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Corral kills every process an AI agent or CI job starts when the timer runs out

An open-source Linux runner uses cgroups or a subreaper plus pidfds to guarantee no child process outlives a timed-out command, exiting with code 120 if it cannot prove the tree is empty.

Corral kills every process an AI agent or CI job starts when the timer runs out

Corral, a new open-source runner that surfaced on Hacker News via the Cardinal44/corral repository, tackles an unglamorous but persistent automation problem: commands that time out while leaving stray processes behind. The tool executes a command under a deadline and promises that, by the time it returns, every process the command spawned is dead. It verifies that claim itself, and if it cannot prove the tree is empty it exits with status code 120.

The failure modes it targets

According to the project's README, AI coding agents and CI pipelines routinely execute shell commands they did not author. Most runners signal the child process or its process group, then wait for the output pipes to close. That model breaks in three common situations. A daemon that forks twice and calls setsid() lands in a different process group and session with init as its parent, so the signal never reaches it. A background process holding stdout or stderr open means the runner never sees end-of-file and hangs after the command exits. And a process that ignores SIGTERM simply keeps running. The leftovers tie up ports, file locks and CPU, and they can make the next run fail.

Two cleanup strategies

Corral manages the entire process tree rather than only its direct child, in two ways. In enforced mode the command runs inside its own cgroup v2 group; the kernel keeps every descendant in that group, and one write to the cgroup.kill file terminates all of them at once. This mode also enforces memory and process-count limits and requires a delegated cgroup, which a bundled script provisions using systemd-run. Fallback mode needs no cgroup: corral acts as a child subreaper so orphaned processes are reparented to it, then scans /proc by parent, process group and session until nothing remains, signalling stragglers through pidfds so a recycled PID is never signalled by mistake. In both modes the run ends when the command exits, not when its pipes close, and after each run corral stops and reaps processes until none are left; in enforced mode the kernel must also report the group empty and the group directory must be removable. The default setting picks enforced mode when it is available.

Configuration and exit codes

Options include a wall-clock deadline, the grace period between SIGTERM and SIGKILL (2 seconds by default), a verification timeout, a combined cap on stdout and stderr, stdin policy, and machine-readable audit records written with a JSON flag. Those records capture the mode, the limits, why the run ended, per-step timings, and the PIDs of anything that did not stop. The exit codes are conventional where possible — 124 for a timeout, 121 for the memory limit, 122 for the output limit, 128 plus the signal number for signalled commands — with code 120 reserved for an unprovable cleanup, overriding every other code. Enforced mode needs Linux 5.14 or later (5.11 minimum otherwise), prebuilt binaries target x86-64 with glibc 2.36, and the project is MIT-licensed. The documentation is explicit that this is not a security sandbox: file access, network access and privileges are not restricted.

What the benchmarks show

The README ships a comparison harness that runs ten fault-injection programs — daemons, SIGTERM-ignorers, mass forkers, pipe holders — against corral in both modes, a timeout-style runner, and a naive Python runner that kills only its direct child, each with a 2-second limit. Corral left zero survivors in every test in both modes. The timeout runner left one process alive in the 200-children test in one of five runs (the author notes it was uutils coreutils 0.8.0 rather than GNU's), while the Python runner hung outright on several tests and left as many as 192 processes running. Enforced mode also stopped a test allocating 256 MiB at a 64 MiB limit in 0.12 seconds. Overhead is modest: running /bin/true took a median of about 11.3 ms in enforced mode and 6.6 ms in fallback, versus 0.83 ms without corral, and a kill was verified within roughly 9–12 ms of the deadline.

Stated limitations

The documentation is candid about gaps. Work the command starts outside its own process tree — via systemd-run, D-Bus or at — is invisible to corral. In enforced mode, a process that writes to cgroupfs itself can leave the group. In fallback mode a setuid child cannot be signalled, which triggers exit code 120. If corral itself is killed with SIGKILL, only its direct child is guaranteed to stop. There is no PTY support, no CPU-time limit, and it runs only on Linux.

Why it matters

Leftover daemons are a classic source of flaky CI and, increasingly, of misbehaving AI agents firing off shell commands with no human watching. Corral turns cleanup from a best-effort hope into a checkable contract: the process tree is provably empty, or you find out through a distinct exit code and an audit record. It complements rather than replaces sandboxing, but it closes a gap that most runners, including the standard timeout utility, only partially address.

  • #process-management
  • #ci-cd
  • #ai-agents
  • #linux
  • #open-source

Related posts