The runners from my ARC post live on an AKS cluster with node auto provisioning (NAP), which adds their nodes when jobs queue up and deletes them again afterwards. NAP is Azure’s managed Karpenter. I moved the runners onto it at the start of 2026 and let it pick spot VMs for the first half of the year, until I counted the jobs that went down with those VMs.
Short version: give the runners a NodePool of their own with on-demand VMs that deletes empty nodes after a short wait (a minute, in my case), and have a job-started hook put karpenter.sh/do-not-disrupt on every runner Pod that picks up a job. On spot, roughly one job in 300 died, and its log said nothing except that the runner had lost its connection.
Moving the cluster from the Cluster Autoscaler to NAP is a story for another post. The numbers here come from the cluster’s metrics, Karpenter’s logs and GitHub’s job records between January and October 2026.
How Karpenter differs from the Cluster Autoscaler
The Cluster Autoscaler needs its node pools defined up front. Every pool is a VM scale set with a single VM size. When Pods can’t be scheduled, the autoscaler adds VMs to a pool, and it takes them away once nodes have sat underused for a while. My runners used to have a pool of Standard_D8ds_v6 VMs, split into one pool per availability zone because the autoscaler doesn’t know about zones.
Karpenter has no fixed pools. Whenever the scheduler can’t place a Pod, Karpenter picks a fitting VM from everything a NodePool allows and creates it directly. Removing nodes is its job as well: empty ones go after a wait you set, and a node whose Pods would fit on fewer or cheaper VMs gets replaced, which Karpenter calls consolidation. With NAP, Karpenter runs inside the managed AKS control plane, and the NodePools and the AKSNodeClass for the node image and OS disk are the only parts I write myself. Microsoft has a side-by-side comparison if you want the long version.
Give the runners their own NodePool
A NodePool tells Karpenter which VMs it may create and how it gets rid of them again. Here’s the one for my runners, cut down to what matters for this post:
apiVersion: karpenter.sh/v1kind: NodePoolmetadata: name: gha-runnerspec: template: metadata: labels: purpose: github-actions spec: nodeClassRef: group: karpenter.azure.com kind: AKSNodeClass name: default requirements: - key: kubernetes.io/arch operator: In values: ["amd64"] - key: karpenter.sh/capacity-type operator: In values: ["on-demand"] - key: karpenter.azure.com/sku-family operator: In values: ["D", "E"] - key: karpenter.azure.com/sku-cpu operator: Gt values: ["4"] - key: karpenter.azure.com/sku-memory operator: Gt values: ["16383"] taints: - key: github-actions value: "true" effect: NoSchedule terminationGracePeriod: 24h disruption: consolidationPolicy: WhenEmptyOrUnderutilized consolidateAfter: 1m budgets: - nodes: "100%" reasons: ["Empty"] - nodes: "10%" reasons: ["Underutilized", "Drifted"]The VM size is up to Karpenter. Anything from the D or E series with more than 4 vCPUs and at least 16 GiB of memory qualifies (sku-memory counts in MiB, hence the odd number). My full NodePool also pins the VM generation and asks for room for an ephemeral OS disk. Out of everything that qualifies, Karpenter takes the cheapest VM that fits the waiting Pods, which in one week in September meant a Standard_D8alds_v6 two times out of three. Bursts get bigger machines. One evening it started a single Standard_E20ads_v6 and packed 19 runners onto it.
The taint keeps other workloads off these nodes. The runner scale set tolerates it and selects the label, and each runner Pod requests 1 CPU and 4 GiB of memory. Those requests have to be there, because Karpenter sizes nodes by them:
template: spec: nodeSelector: purpose: github-actions tolerations: - key: github-actions operator: Equal value: "true" effect: NoSchedule resources: requests: cpu: "1" memory: 4Gi limits: memory: 16Gi containers: - name: runner image: ghcr.io/nobbs/actions-runner:<VERSION> command: ["/home/runner/run.sh"]I put the resources on the Pod as a whole instead of per container. Pod-level resources have been beta and on by default since Kubernetes 1.34, and Karpenter counts them the same way it counts container requests.
Keep Karpenter away from running jobs
A runner halfway through a job looks like any other Pod to Karpenter, and if consolidation decides its node should go, the job goes with it. Karpenter’s disruption docs list an annotation for Pods that have to stay where they are. It can:
block Karpenter from voluntarily disrupting and draining pods
Setting it in the runner template would protect the idle runners too, and with a few of them always waiting, their nodes would never be removed. I want it on busy runners only. A job-started hook handles that. The runner calls the hook just before a job begins, and the hook annotates the runner’s own Pod.
For that, the hook needs kubectl. My runner image has it, and the service account needs permission to patch Pods in the runner namespace. ARC runs runner Pods as the service account <SCALE_SET>-gha-rs-no-permission, where the scale set’s name defaults to the Helm release name. With the scale set from the ARC post, that makes it arc-runner-set-gha-rs-no-permission:
apiVersion: rbac.authorization.k8s.io/v1kind: Rolemetadata: name: pod-annotator namespace: arc-runnersrules: - apiGroups: [""] resources: ["pods"] verbs: ["get", "patch"]---apiVersion: rbac.authorization.k8s.io/v1kind: RoleBindingmetadata: name: pod-annotator namespace: arc-runnerssubjects: - kind: ServiceAccount name: arc-runner-set-gha-rs-no-permission namespace: arc-runnersroleRef: kind: Role name: pod-annotator apiGroup: rbac.authorization.k8s.io---apiVersion: v1kind: ConfigMapmetadata: name: runner-hooks namespace: arc-runnersdata: job-started.sh: | #!/bin/sh KUBE_DIR="/home/runner/.kube-api" kubectl \ --server="https://kubernetes.default.svc" \ --certificate-authority="${KUBE_DIR}/ca.crt" \ --token="$(cat "${KUBE_DIR}/token")" \ --namespace="$(cat "${KUBE_DIR}/namespace")" \ annotate pod "${HOSTNAME}" karpenter.sh/do-not-disrupt="true" --overwrite > /dev/null 2>&1 || trueSince the Pod name doubles as the container’s hostname, ${HOSTNAME} tells the hook which Pod to annotate, and the scale set’s values mount the hook and point ACTIONS_RUNNER_HOOK_JOB_STARTED at it. I also switch off the automatic service account token for the Pod and mount a token into the runner container alone:
template: spec: automountServiceAccountToken: false containers: - name: runner image: ghcr.io/nobbs/actions-runner:<VERSION> command: ["/home/runner/run.sh"] env: - name: ACTIONS_RUNNER_HOOK_JOB_STARTED value: /hooks/job-started.sh volumeMounts: - name: hooks mountPath: /hooks - name: kube-api mountPath: /home/runner/.kube-api readOnly: true volumes: - name: hooks configMap: name: runner-hooks - name: kube-api projected: sources: - serviceAccountToken: path: token expirationSeconds: 3600 - configMap: name: kube-root-ca.crt items: - key: ca.crt path: ca.crt - downwardAPI: items: - path: namespace fieldRef: fieldPath: metadata.namespaceJobs run in that same container. A workflow can read the token and do whatever the hook can do, which is why the Role allows nothing beyond reading and patching Pods in the runner namespace. An admission policy can lock it down further by rejecting any change from the runner’s service account that touches more than the karpenter.sh/do-not-disrupt annotation. Kyverno can do that, and so can a built-in ValidatingAdmissionPolicy.
Over one week in September, Karpenter moved hundreds of idle runners to other nodes, and ARC recreated them there. It backed off from nodes with a busy runner. Its events give the reason as Pod has "karpenter.sh/do-not-disrupt" annotation.
Count what spot costs
Until the start of July, the runner NodePool accepted spot VMs as well as on-demand ones. Karpenter goes for the cheaper offer. On the last spot morning I have logs for, every runner node it created was spot. In West Europe, a Standard_D8alds_v6 costs $0.461 an hour on-demand and around $0.085 as spot, which brought June down to $0.008 per job, where the same nodes on-demand would have cost about $0.04.
Azure can take a spot VM back whenever it needs the capacity, with up to 30 seconds of warning. Back then, Karpenter on AKS did nothing with that warning, so the VM vanished and took the runner Pod with it. GitHub then waits for the runner to check in again, gives up after roughly nine minutes, and fails the job with this message:
The self-hosted runner lost communication with the server.Spot appears nowhere in it, and ARC’s metrics showed no spike, so I went digging. When Azure announces an eviction, AKS sets a VMEventScheduled condition on the node, and the cluster’s metrics record which runner Pods sat on that node when it disappeared. GitHub’s job records carry the runner’s name, which is also the Pod’s name, so matching the two is mostly a join on one column.
In June, roughly one runner node in five got evicted. About one job in 300 was running on a node when it went and died with it, on almost every weekday. Nearly all of those jobs hung for eight minutes or more before they failed. The worst one threw away more than two hours of work, and a few were already re-runs, one of them on its seventh attempt.
At the start of July I moved the runners to on-demand because of the silent failures. A job that hangs for nine minutes and then blames the network gives nobody a clue that a spot VM disappeared underneath it. A little later, AKS added an in-VM spot eviction signal to NAP, and I want to give spot another try with that in place.
Remove empty nodes fast
On-demand VMs cost about five times as much per hour, and every idle runner node got expensive. Until July, Karpenter waited 15 minutes before removing an empty one. Over the summer I took that down to a minute. That wait is consolidateAfter in the NodePool above, the time Karpenter gives a node to pick up new work before it considers removing it.
The budgets let all empty nodes go at once and replace underused or drifted nodes, after a new node image for example, at most 10% at a time.
Between June and September, node time per job fell to about a third. Half of the runner nodes are gone again within ten minutes of starting. A job now costs a little over a cent, around one and a half times what it cost on spot, on VMs that are five times pricier per hour.
Karpenter starts a node for every waiting runner, and by the time that node is up, some runners have already landed on a node that just finished a job. Roughly one runner node in six never runs a single job.
Keep a few runners warm during business hours
A job that finds no idle runner waits for a new node, and the fast cleanup makes that more common. Getting a new node to accept Pods takes two minutes or so after Karpenter decides to create it, and pulling the runner image, close to a gigabyte, adds another 20 seconds. With an idle runner waiting, a job starts right away.
A cronjob keeps five runners idle during business hours and one at all other times. During business hours, hardly any job waits long. Outside them, up to one job in five waits more than a minute, because once a job grabs the single warm runner, its replacement needs a new node.
What a job costs now
In January, still on the Cluster Autoscaler and on-demand VMs, a runner job cost $0.150 in VM time, at West Europe list prices without disks or discounts. By September it was $0.013, mostly because of load. In March, with roughly as many jobs as in January, NAP used about as many node-hours as the Cluster Autoscaler had. Between January and September the number of jobs grew many times over and all those jobs share the same idle capacity, with the business-hours warm pool and the one-minute cleanup doing the rest.
References
- Node auto provisioning in AKSlearn.microsoft.com
- Migrate from the Cluster Autoscaler to NAPlearn.microsoft.comwith Microsoft’s comparison of the two
- NAP NodePoolslearn.microsoft.com and AKSNodeClass
- Karpenter disruptionkarpenter.sh (consolidation, budgets and
do-not-disrupt) - Run scripts before or after a jobdocs.github.comthe job-started hook
- Azure Spot Virtual Machineslearn.microsoft.com and Scheduled Events