Skip to content

Why user namespaces can't run Docker-in-Docker yet

Setting hostUsers: false on a Kubernetes Pod maps container root to nobody, but dockerd still can't start a container. Here's the exact wall, and the KEP that should remove it.

My ARC runners run Docker-in-Docker as a privileged sidecar, which is root on the node. The obvious fix is a user namespace, GA since Kubernetes 1.36: root in the Pod becomes nobody on the node. I couldn’t get Docker-in-Docker to run under it.

Short version: hostUsers: false gets you a working Docker daemon and image storage, but not a running container: every variation dies at the same place, creating the child cgroup for the container Docker is about to start, and the Pod spec can’t change that yet. KEP-5474 fixes it, but it’s targeted at alpha in 1.38: off by default, behind a feature gate, and not something to move production jobs onto.

The one line, and where it lands you

The knob is this:

hostUsers: false

With it, the daemon comes up fine. On a stock AKS node (Kubernetes 1.36.3, Ubuntu 24.04, kernel 6.8, containerd 2.3.3) the uid_map shows a real user namespace, docker info works, and images pull and unpack (docker:29.6.0-dind). Overlayfs inside a user namespace is no problem. Then the first container fails:

Terminal window
$ docker run alpine true
... runc create failed: unable to apply cgroup configuration: mkdir /sys/fs/cgroup/docker: permission denied

Inside the daemon container, the cgroup filesystem is mounted read-write but owned by nobody:

Terminal window
$ stat -c '%U %a' /sys/fs/cgroup
nobody 555

The Pod’s cgroup directory belongs to host root, and host root doesn’t exist inside the user namespace, so it shows up as nobody. It isn’t delegated to the container’s root, so runc can’t create the child cgroup it needs. Starting a daemon and starting a container turn out to be different problems, and only the first one works.

Every variation hits the same wall

I tried the obvious ways around it, and they all fail at cgroups or before:

  • Keep privileged: true and add hostUsers: false (the “privileged, but not really” trick from BuildKit’s examples and CERN). Daemon up, images pull, docker run fails on mkdir /sys/fs/cgroup/docker: permission denied, cgroupfs is read-write but nobody-owned.
  • Drop privileged, add capabilities (SYS_ADMIN, unconfined seccomp and AppArmor). Same wall from the other side: cgroupfs is read-only, so the same mkdir fails with read-only file system.
  • Rootless (docker:dind-rootless, no privileged). Doesn’t even reach the daemon. rootlesskit can’t fork inside the nested user namespace and fails with fork/exec /proc/self/exe: operation not permitted.

The missing piece is the same in all three: cgroup delegation, handing the Pod’s cgroup subtree to the container’s root so it can create children below it. The kubelet and containerd decide that when they set the container up, and no securityContext field reaches it. Adding capabilities or loosening seccomp changes nothing, because the failure isn’t a capability check: it’s a read-only mount in one case and file ownership in the other.

The fix that’s coming

KEP-5474 “Enable Writable Cgroups” is the direct fix. It adds a field:

securityContext:
cgroupOptions:
mountMode: Writable

On cgroup v2 this mounts /sys/fs/cgroup read-write for the container’s own subtree via the kernel’s nsdelegate, so a nested runtime can create child cgroups but can’t reach up and change the Pod’s limits. Docker-in-Docker is the KEP’s first user story. It’s targeted at alpha in 1.38, behind a CgroupOptions feature gate, and needs matching cgroupWritable support in containerd underneath.

Two limits. My tests died at the cgroup wall, daemon up and images pulled, so nothing past it is tested. And cgroups are one piece of what a nested Docker needs, next to /proc, /sys, mount and seccomp handling, which this KEP doesn’t touch.

References

Published
Reading
3 min