My ARC runners run Docker-in-Docker as a privileged sidecar, which is root on the node. The obvious fix is a user namespace, GA since Kubernetes 1.36: root in the Pod becomes nobody on the node. I couldn’t get Docker-in-Docker to run under it.
Short version: hostUsers: false gets you a working Docker daemon and image storage, but not a running container: every variation dies at the same place, creating the child cgroup for the container Docker is about to start, and the Pod spec can’t change that yet. KEP-5474 fixes it, but it’s targeted at alpha in 1.38: off by default, behind a feature gate, and not something to move production jobs onto.
The one line, and where it lands you
The knob is this:
hostUsers: falseWith it, the daemon comes up fine. On a stock AKS node (Kubernetes 1.36.3, Ubuntu 24.04, kernel 6.8, containerd 2.3.3) the uid_map shows a real user namespace, docker info works, and images pull and unpack (docker:29.6.0-dind). Overlayfs inside a user namespace is no problem. Then the first container fails:
$ docker run alpine true... runc create failed: unable to apply cgroup configuration: mkdir /sys/fs/cgroup/docker: permission deniedInside the daemon container, the cgroup filesystem is mounted read-write but owned by nobody:
$ stat -c '%U %a' /sys/fs/cgroupnobody 555The Pod’s cgroup directory belongs to host root, and host root doesn’t exist inside the user namespace, so it shows up as nobody. It isn’t delegated to the container’s root, so runc can’t create the child cgroup it needs. Starting a daemon and starting a container turn out to be different problems, and only the first one works.
Every variation hits the same wall
I tried the obvious ways around it, and they all fail at cgroups or before:
- Keep
privileged: trueand addhostUsers: false(the “privileged, but not really” trick from BuildKit’s examples and CERN). Daemon up, images pull,docker runfails onmkdir /sys/fs/cgroup/docker: permission denied, cgroupfs is read-write butnobody-owned. - Drop privileged, add capabilities (
SYS_ADMIN, unconfined seccomp and AppArmor). Same wall from the other side: cgroupfs is read-only, so the samemkdirfails withread-only file system. - Rootless (
docker:dind-rootless, no privileged). Doesn’t even reach the daemon. rootlesskit can’t fork inside the nested user namespace and fails withfork/exec /proc/self/exe: operation not permitted.
The missing piece is the same in all three: cgroup delegation, handing the Pod’s cgroup subtree to the container’s root so it can create children below it. The kubelet and containerd decide that when they set the container up, and no securityContext field reaches it. Adding capabilities or loosening seccomp changes nothing, because the failure isn’t a capability check: it’s a read-only mount in one case and file ownership in the other.
The fix that’s coming
KEP-5474 “Enable Writable Cgroups” is the direct fix. It adds a field:
securityContext: cgroupOptions: mountMode: WritableOn cgroup v2 this mounts /sys/fs/cgroup read-write for the container’s own subtree via the kernel’s nsdelegate, so a nested runtime can create child cgroups but can’t reach up and change the Pod’s limits. Docker-in-Docker is the KEP’s first user story. It’s targeted at alpha in 1.38, behind a CgroupOptions feature gate, and needs matching cgroupWritable support in containerd underneath.
Two limits. My tests died at the cgroup wall, daemon up and images pulled, so nothing past it is tested. And cgroups are one piece of what a nested Docker needs, next to /proc, /sys, mount and seccomp handling, which this KEP doesn’t touch.
References
- Kubernetes user namespaceskubernetes.io
- KEP-5474: Enable Writable Cgroupsgithub.com
- Rootless container builds on Kuberneteskubernetes.web.cern.chCERN, the “privileged, but not really” pattern
- ARC issue #4189github.comothers reporting the same failures
- lhns/k0s-flux-examplegithub.comsame cgroup wall on k0s, kernel 6.12