deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Rootless Rust container engine boxr explains its namespace and uid mapping design

A dev.to post from the boxr project explains how the Rust OCI engine isolates containers without root, using a single-threaded trampoline process and uid maps written by the parent.

Rootless Rust container engine boxr explains its namespace and uid mapping design

boxr, a rootless OCI container engine written in Rust, has published a walkthrough of how it isolates containers without a daemon, sudo, or a setuid helper. The post on dev.to traces every Linux step between invoking boxr run and the container process starting. Its organizing principle: enter a user namespace first, then carry out all remaining setup from inside it.

Why threads block the obvious approach

According to the dev.to post, the boxr CLI is a multi-threaded tokio application, and the Linux kernel rejects unshare(CLONE_NEWUSER) from any process that already has threads, returning EINVAL. Because tokio spawns its worker threads before application code runs, the CLI itself can never create the user namespace.

The fix is a re-exec. Before the async runtime starts, boxr launches itself again as an internal trampoline subprocess. That process is single-threaded, so it can unshare the user namespace cleanly, and all namespace setup runs there as straightforward sequential code before any async runtime exists. The author argues this sidesteps an entire class of problems that come from mixing namespace setup with an async runtime.

A two-process handshake for uid mapping

Once a child unshares the user namespace it is root inside that namespace, but without a uid mapping it cannot accomplish anything, and only the parent is permitted to write the maps. So the trampoline forks: the child unshares CLONE_NEWUSER, and the two processes synchronize over a Unix socketpair with a simple ready/done exchange while the parent writes /proc/[pid]/uid_map.

Two mapping paths exist, per the post. If newuidmap and newgidmap are installed and /etc/subuid and /etc/subgid define ranges for the user, those tools are used: container uid 0 maps to the host user, and container uids from 1 upward map onto the subordinate range. Otherwise boxr writes the maps directly — a single entry mapping uid 0 to the host uid, written only after sending "deny" to setgroups, a step the kernel mandates before accepting a gid_map.

Network first, then the remaining namespaces

With root-in-namespace secured, the child next unshares the network namespace, deliberately ahead of the others, so the parent can attach a pasta process, or the pure-Rust usernet TAP engine can start its worker, before the namespace topology grows more complicated.

The child then unshares the PID, mount, UTS and IPC namespaces. The cgroup namespace is opt-in via annotation, and IPC and UTS can likewise be left on the host. A second fork makes the grandchild PID 1 in the new PID namespace. From there the engine bind-mounts the rootfs onto itself (pivot_root requires a mountpoint), calls pivot_root with a chroot fallback, and execs the container init. The author emphasizes that no step in the chain executes with root privileges on the host.

The limits the project acknowledges

The post is candid about what the design does not deliver. Without subuid/subgid configured, the container sees a single mapped uid, which suits development containers but falls short of a full multi-user mapping. Rootless networking goes through pasta or the user-mode TAP engine rather than a veth pair, so raw sockets and certain packet types behave differently. And the described path is Linux-only; the macOS and Windows runtimes take entirely different routes through Virtualization.framework and WSL2.

Interested readers can find the trampoline and namespace ordering in src/runtime/linux.rs and the uid/gid mapping strategy in src/security/mod.rs within the kchaitanya863/boxr repository.

Why it matters

Rootless containers remove one of the sharpest edges of container usage: dependence on a privileged daemon. The post works as both an implementation note and a compact explanation of mechanics most container users never see — why uid maps exist at all, why setgroups must be denied before a gid_map is accepted, why pivot_root demands a mountpoint. It also documents a constraint that any async runtime wanting namespaces must solve, namely the kernel's refusal to unshare from threaded processes, and shows re-exec as a clean answer. The frankness about the single-uid fallback and the networking differences is the kind of detail that lets prospective users judge whether a young runtime fits their workload.

  • #rust
  • #containers
  • #linux
  • #namespaces
  • #open-source

Related posts