Skip to content

Run self-hosted GitHub Actions runners with ARC

How I run self-hosted GitHub Actions runners inside private Azure networks: Actions Runner Controller, a GitHub App, a small runner image and one test job.

The Azure infrastructure I work on uses private networking only, so GitHub’s hosted runners can’t reach its clusters. Self-hosted runners can, because they run inside those networks. They also cost nothing beyond the Azure resources underneath, at least for now: GitHub announced a $0.002 per minute charge for self-hosted runners from March 2026, then postponed it to re-evaluate.

Mine live on dedicated node pools in a dev AKS cluster, where Node Auto Provisioning (NAP), Azure’s managed Karpenter, adds and removes nodes as jobs come and go. I’ve run setups like this for about three years, this one for the last 12 months, and last month it ran about 50,000 jobs.

Short version: install the Actions Runner Controller once per cluster, give it a GitHub App, then deploy a runner scale set per group of runners. Each job gets a fresh Pod built from a small image that contains the Actions runner and the few tools every workflow needs.

Tested with ARC chart 0.14.2, runner 2.337.0, kubectl 1.37.1 and Helm 4.3, on a local kind cluster.

Know the moving parts

Actions Runner Controller (ARC) is a Kubernetes operator. It comes as two Helm charts:

  • gha-runner-scale-set-controller installs the controller and its CRDs. You install it once per cluster.
  • gha-runner-scale-set configures one runner scale set: which GitHub organization or repository it registers with, how many runners it may start, and what the runner Pod looks like. You install it once per scale set, and they all share the controller.

Avoid the legacy ARC docs

ARC started in 2020 as a community project that GitHub later took over. In 2023 GitHub added the runner scale set mode these two charts install, and made it generally available that June.

The older modes are still in the repository, but they’re a different controller: the actions-runner-controller chart installs it, and it manages its own CRDs like RunnerDeployment and HorizontalRunnerAutoscaler in the actions.summerwind.net API group. The ARC README calls them legacy, maintained by the community only.

Plenty of blog posts and answers online still describe the legacy modes. If an example uses actions.summerwind.net resources or the actions-runner-controller chart, it’s not for the setup in this post.

Follow a job through ARC

For each scale set, the controller starts a listener Pod in the controller’s namespace, not next to the runners. The listener holds a long-poll connection to GitHub and waits for jobs. When one is queued for its scale set, it raises the desired runner count, and the controller creates an ephemeral runner Pod. The runner registers, takes exactly one job, and exits. The controller then deletes the Pod.

long-poll for jobs job queued raise runner count create Pod register, run one job delete Pod when done GitHub Listener Controller Runner Pod

Every connection starts inside the cluster. The runners need outbound access to GitHub, but nothing has to reach in.

Gather the requirements

  • A Kubernetes cluster with kubectl access and sufficient permissions to install CRDs, and a way to install Helm charts: Helm 3.8 or later, or a GitOps tool like Argo CD or Flux. The charts come from an OCI registry, and Helm only supports those by default since 3.8.
  • A GitHub organization or repository for the runners, and a GitHub App to register them. The next section covers the App.
  • A container registry the runner image is stored in, reachable from the cluster.

Allow outbound traffic to GitHub

The controller, listener and runner Pods all need outbound HTTPS to GitHub, and that means more than github.com. GitHub’s self-hosted runners reference lists the domains by function. These are the ones every setup needs:

Needed for Domains
Essential operations github.com, api.github.com, *.actions.githubusercontent.com
Downloading actions codeload.github.com
Logs, artifacts and caches results-receiver.actions.githubusercontent.com, *.blob.core.windows.net

Runner updates, GitHub Packages, Git LFS and release assets each add a few more, and then there’s whatever your workflows download. Note that *.blob.core.windows.net matches every Azure storage account, not just GitHub’s. If your cluster runs Cilium, CiliumCIDRGroup keeps the IP-based part of those egress rules readable.

Keep the logs

Runner Pods are deleted right after their job, and their container logs go with them.

Those logs don’t include the job output. What you see in the Actions tab is sent to GitHub, not printed to stdout, so if you need it anywhere else, you have to scrape it from the runner Pod while the job runs.

The container logs hold the runner’s own diagnostics, next to the controller and listener logs. They’re what you need when a runner fails to start or register, and GitHub’s setup guide recommends collecting them before you go to production:

While it is not required for ARC to be deployed, we recommend ensuring you have implemented a way to collect and retain logs from the controller, listeners, and ephemeral runners before deploying ARC in production workflows.

Create the GitHub App

ARC can also use a personal access token. That might be fine for personal use, but it has no place in an enterprise setup. A token always belongs to one user account, so your runners depend on that person’s account and access. A GitHub App can be owned by the organization itself, and the tokens it issues are short-lived.

GitHub has step-by-step instructions. For a scale set at organization level, the App needs these permissions:

  • Repository permissions: Metadata: read-only.
  • Organization permissions: Self-hosted runners: read and write.

A scale set for a single repository needs Administration: read and write on the repository instead. Install the App on the organization, and note three things: the App ID, the installation ID from the installation page’s URL, and the private key .pem file you generate.

Install the controller

The controller gets its own namespace:

Terminal window
helm install arc \
--namespace arc-systems --create-namespace \
--version 0.14.2 \
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set-controller

Check that the controller runs and the CRDs exist:

Terminal window
$ kubectl -n arc-systems get pods
NAME READY STATUS RESTARTS AGE
arc-gha-rs-controller-674cf6b948-497xb 1/1 Running 0 10s
$ kubectl get crd | grep actions.github.com
autoscalinglisteners.actions.github.com Namespaced v1alpha1(storage) 2026-09-24T19:04:21Z
autoscalingrunnersets.actions.github.com Namespaced v1alpha1(storage) 2026-09-24T19:04:21Z
ephemeralrunners.actions.github.com Namespaced v1alpha1(storage) 2026-09-24T19:04:22Z
ephemeralrunnersets.actions.github.com Namespaced v1alpha1(storage) 2026-09-24T19:04:22Z

Pin --version, and use the same version for the scale sets later.

That’s awkward by hand. A GitOps tool like Argo CD makes it easier, because it renders the chart and applies the CRDs like any other manifest, so they get updated together with the controller.

One catch: ARC’s CRDs are big, between 300 and 600 KB each. That’s over the 256 KB limit for the annotation a client-side kubectl apply stores, so turn on the ServerSideApply=true sync option for them.

Build a small runner image

GitHub publishes a runner image, ghcr.io/actions/actions-runner, but it’s very bare-bones: the runner, the Docker CLI, Git and a few basics like curl and jq. That’s a long way from the hundreds of tools on GitHub-hosted runners.

I build my own from Ubuntu 24.04 for two reasons. Jobs get the same tool versions every time, and common tools like kubectl, Helm and the Azure CLI aren’t downloaded again in every job.

The cost is maintenance. Every tool in the image is another thing to update, scan and test, and it makes the image bigger for everyone. My rule is that a tool goes into the image only if most workflows need it. Everything else stays in the workflow, installed by an action.

Renovate and one small test per tool take most of the pain out of the upkeep. Renovate opens a pull request whenever a pinned version has a new release, and a small Bats test per tool, run in CI after every build, catches the updates that break something. My tests carry the same # renovate: comments as the Dockerfile below, so one pull request bumps both the tool and the version its test expects.

Here is a cut-down version of my image, with kubectl as the one extra tool:

Dockerfile
FROM ubuntu:24.04
# renovate: datasource=github-releases depName=runner packageName=actions/runner
ARG RUNNER_VERSION=2.337.0
# renovate: datasource=github-releases depName=kubectl packageName=kubernetes/kubernetes
ARG KUBECTL_VERSION=1.37.1
ENV DEBIAN_FRONTEND=noninteractive
ENV RUNNER_MANUALLY_TRAP_SIG=1
ENV ACTIONS_RUNNER_PRINT_LOG_TO_STDOUT=1
RUN apt-get update \
&& apt-get install -y --no-install-recommends ca-certificates curl git jq sudo \
&& useradd --create-home --uid 1001 --shell /bin/bash runner \
&& echo "runner ALL=(ALL) NOPASSWD:ALL" > /etc/sudoers.d/runner \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /home/runner
RUN curl -fLo runner.tar.gz "https://github.com/actions/runner/releases/download/v${RUNNER_VERSION}/actions-runner-linux-x64-${RUNNER_VERSION}.tar.gz" \
&& sha=$(curl -fsSL "https://api.github.com/repos/actions/runner/releases/tags/v${RUNNER_VERSION}" \
| jq -r .body | grep "BEGIN SHA linux-x64" | grep -oE '[a-f0-9]{64}') \
&& echo "${sha} runner.tar.gz" | sha256sum -c - \
&& tar xzf runner.tar.gz && rm runner.tar.gz \
&& ./bin/installdependencies.sh \
&& chown -R runner:runner /home/runner \
&& rm -rf /var/lib/apt/lists/*
RUN curl -fLo /usr/local/bin/kubectl "https://dl.k8s.io/release/v${KUBECTL_VERSION}/bin/linux/amd64/kubectl" \
&& echo "$(curl -fsSL "https://dl.k8s.io/release/v${KUBECTL_VERSION}/bin/linux/amd64/kubectl.sha256") /usr/local/bin/kubectl" \
| sha256sum -c - \
&& chmod 755 /usr/local/bin/kubectl
USER runner

A few lines need an explanation:

  • The runner’s release notes list a SHA-256 checksum for every archive. The build reads it from the GitHub API and refuses to continue if the download doesn’t match. kubectl gets the same treatment, with the .sha256 file published next to the binary.
  • The checksum lookup is an unauthenticated API call, which GitHub limits to 60 requests an hour per IP address. One build needs one, but many builds behind a shared NAT address can run out.
  • installdependencies.sh ships with the runner. It installs the libraries the runner’s .NET runtime needs, such as ICU.
  • The runner runs as the unprivileged user runner, with passwordless sudo. That’s what GitHub’s own image does too, and plenty of marketplace actions call sudo apt-get.
  • RUNNER_MANUALLY_TRAP_SIG=1 makes run.sh pass SIGTERM and SIGINT on to the runner process, so the runner can shut down cleanly when its Pod is deleted. ACTIONS_RUNNER_PRINT_LOG_TO_STDOUT=1 sends the runner’s own diagnostic logs to stdout, where your log collector can find them.

The image comes out at 1.4 GB. About 680 MB of that is the runner itself, mostly the Node.js versions it bundles for JavaScript actions.

Build and push it:

Terminal window
docker build --platform linux/amd64 --tag ghcr.io/nobbs/actions-runner:<VERSION> .
docker push ghcr.io/nobbs/actions-runner:<VERSION>

Deploy one runner scale set

The runners get a namespace of their own, and the GitHub App secret goes there too:

Terminal window
kubectl create namespace arc-runners
kubectl create secret generic arc-github-app \
--namespace arc-runners \
--from-literal=github_app_id=<APP_ID> \
--from-literal=github_app_installation_id=<INSTALLATION_ID> \
--from-file=github_app_private_key=<PATH_TO_PEM>

The scale set’s values.yaml only needs a few keys:

values.yaml
githubConfigUrl: https://github.com/<ORG>
githubConfigSecret: arc-github-app
minRunners: 0
maxRunners: 3
template:
spec:
containers:
- name: runner
image: ghcr.io/nobbs/actions-runner:<VERSION>
command: ["/home/runner/run.sh"]

Don’t rename the container. The deployment guide is strict about it:

The runner container must be named runner. Otherwise, it will not be configured properly to connect to GitHub.

minRunners: 0 means no idle Pods, and each job waits for a Pod to start. Raise it if that wait matters more to you than the idle capacity. I keep 5 runners idle during business hours. The chart has no schedule for that, so a CronJob patches minRunners up in the morning and back down in the evening.

If your organization uses runner groups, runnerGroup puts the scale set into one.

Install it:

Terminal window
helm install arc-runner-set \
--namespace arc-runners \
--version 0.14.2 \
--values values.yaml \
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set

The Helm release name, arc-runner-set here, is the scale set’s name in GitHub, and runnerScaleSetName in the values overrides it. Workflows can put the name in runs-on, but then renaming the scale set means changing every one of them.

Labels avoid that. The legacy controller had them, the scale set mode shipped without them, and bringing them back was a popular request, with 43 upvotes on the issue. They returned in chart 0.14.0 as scaleSetLabels. In the values:

values.yaml
scaleSetLabels:
- linux
- kubectl

A workflow then asks for all of them as an array:

runs-on: [linux, kubectl]

Run one test job

.github/workflows/arc-test.yaml
name: arc-test
on: workflow_dispatch
jobs:
hello:
runs-on: arc-runner-set
steps:
- run: /home/runner/bin/Runner.Listener --version
- run: kubectl version --client

Watch the runner namespace, then start the workflow from the Actions tab:

Terminal window
kubectl get pods -n arc-runners --watch

A runner Pod appears, runs the job and disappears. In my test on kind, the job started about 6 seconds after the listener asked for a runner, and the whole run took 13 seconds.

The job log shows it ran on the custom image:

2.337.0
Client Version: v1.37.1

That’s a working setup. Running Docker inside jobs will follow in a separate post, including how to do it securely, without giving every job a way to escape to root on the host. So will scaling the nodes underneath with NAP.

References

Published
Reading
9 min