Checks
Controller Version
0.14.2 (same code path on master at 03328aa)
Deployment Method
Helm
Checks
To Reproduce
1. Install gha-runner-scale-set-controller and one gha-runner-scale-set with a `githubConfigSecret` whose token is revoked or expired (any deterministic startup failure works: 401 from the registration-token API, 403 from REST rate-limit exhaustion, 404 RunnerScaleSetNotFoundException).
2. Watch the listener pod in the controller namespace: `kubectl get pods -n <controller-ns> -w`.
3. The pod is created, runs the auth handshake, exits non-zero within a few seconds, is deleted by the controller and recreated immediately. It never reaches a steady failed state and the loop never slows down.
4. With N scale sets installed, all N listeners loop in lock-step. Fixing the secret ends the loop, but while it runs, each iteration spends more GitHub API quota.
Describe the bug
The AutoscalingListener reconciler has no failure accounting and no delay between a listener pod terminating and the next pod being created:
- container
Terminated -> deleteListenerPod immediately, ctrl.Result{}:
|
case cs.State.Terminated != nil: |
|
log.Info( |
|
"Listener pod is terminated", |
|
"namespace", listenerPod.Namespace, |
|
"name", listenerPod.Name, |
|
"reason", cs.State.Terminated.Reason, |
|
"message", cs.State.Terminated.Message, |
|
) |
|
|
|
return ctrl.Result{}, r.deleteListenerPod(ctx, &autoscalingListener, &listenerPod, log) |
- pod
NotFound -> r.Create(desiredPod) immediately:
|
case kerrors.IsNotFound(err): |
|
if err := r.publishRunningListener(&autoscalingListener, false); err != nil { |
|
// If publish fails, URL is incorrect which means the listener pod would never be able to start |
|
return ctrl.Result{}, nil |
|
} |
|
|
|
r.ResourceCache.listenerPod.Delete(&autoscalingListener) |
|
desiredPod, err := r.newScaleSetListenerPod( |
|
&autoscalingListener, |
|
&listenerConfigSecret, |
|
&serviceAccount, |
|
&listenerRole, |
|
&listenerRoleBinding, |
|
metricsConfig, |
|
) |
|
if err != nil { |
|
log.Error(err, "Failed to build listener pod") |
|
return ctrl.Result{}, err |
|
} |
|
|
|
log.Info("Creating listener pod", "namespace", desiredPod.Namespace, "name", desiredPod.Name) |
|
if err := r.Create(ctx, desiredPod); err != nil { |
|
log.Error(err, "Unable to create listener pod", "namespace", desiredPod.Namespace, "name", desiredPod.Name) |
|
return ctrl.Result{}, err |
|
} |
|
default: // error |
|
log.Error(err, "Unable to get listener pod", "namespace", autoscalingListener.Namespace, "name", autoscalingListener.Name) |
deleteListenerPod:
|
func (r *AutoscalingListenerReconciler) deleteListenerPod(ctx context.Context, autoscalingListener *v1alpha1.AutoscalingListener, listenerPod *corev1.Pod, log logr.Logger) error { |
|
if err := r.publishRunningListener(autoscalingListener, false); err != nil { |
|
log.Error(err, "Unable to publish runner listener down metric", "namespace", listenerPod.Namespace, "name", listenerPod.Name) |
|
} |
|
|
|
if listenerPod.DeletionTimestamp.IsZero() { |
|
log.Info("Deleting the listener pod", "namespace", listenerPod.Namespace, "name", listenerPod.Name) |
|
if err := r.Delete(ctx, listenerPod); err != nil && !kerrors.IsNotFound(err) { |
|
log.Error(err, "Unable to delete the listener pod", "namespace", listenerPod.Namespace, "name", listenerPod.Name) |
|
return err |
|
} |
|
} |
|
return nil |
|
} |
The listener pod is built with RestartPolicy: Never (
|
RestartPolicy: corev1.RestartPolicyNever, |
), so the kubelet's CrashLoopBackOff never applies either; the controller is the only thing pacing restarts. Both transitions are driven by
Owns(&corev1.Pod{}) watch events, which enqueue with a plain
Add rather than
AddRateLimited, so the workqueue's per-item exponential backoff does not apply (it only applies to requeues/errors). Neither
--workqueue-rate-limiter setting changes this.
Contrast with the EphemeralRunner reconciler, which records failures in status and backs off 5s -> 10s -> 20s -> 40s -> 80s before recreating a runner pod (#4059):
|
// precompute backoff durations for failed ephemeral runners |
|
// the len(failedRunnerBackoff) must be equal to maxFailures + 1 |
|
var failedRunnerBackoff = []time.Duration{ |
|
0, |
|
5 * time.Second, |
|
10 * time.Second, |
|
20 * time.Second, |
|
40 * time.Second, |
|
80 * time.Second, |
|
} |
|
|
|
const maxFailures = 5 |
|
now := metav1.Now() |
|
lastFailure := ephemeralRunner.Status.LastFailure() |
|
backoffDuration := failedRunnerBackoff[len(ephemeralRunner.Status.Failures)] |
|
nextReconciliation := lastFailure.Add(backoffDuration) |
|
if !lastFailure.IsZero() && now.Before(&metav1.Time{Time: nextReconciliation}) { |
|
requeueAfter := nextReconciliation.Sub(now.Time) |
|
log.Info( |
|
"Backing off the next reconciliation due to failure", |
|
"lastFailure", lastFailure, |
|
"nextReconciliation", nextReconciliation, |
|
"requeueAfter", requeueAfter, |
|
) |
|
return ctrl.Result{ |
|
Requeue: true, |
|
RequeueAfter: requeueAfter, |
|
}, nil |
|
} |
The listener gets no equivalent. The existing test It should re-create pod but persist config secret whenever listener container is terminated (
|
It("It should re-create pod but persist config secret whenever listener container is terminated", func() { |
) pins the current behaviour: recreation is immediate and unconditional, even for exit code 0.
Why it matters: every listener start repeats the full auth handshake against GitHub (installation/registration token, then the Actions admin connection; the admin-connection client itself retries 401/403 up to 4 more times in-process) before exiting. When the failure is deterministic (expired/rotated PAT, exhausted REST rate limit, scale set gone server-side), every listener in the cluster enters a create/crash/delete loop of a few seconds and the loop itself burns the remaining quota. In a cluster with several hundred scale sets on one PAT we saw several hundred listener pods in Error, sustained 403 API rate limit exceeded, and kubectl get pods often showed no listener pod at all because each one lived 1-3 s. Recovery after fixing the credential was delayed because the loop had kept the limit pinned. #3942 asked for in-process retries and was closed with "the listener would try to restart and recover on its own" -- that restart path is the intended recovery mechanism, so it should be paced.
Describe the expected behavior
A listener pod that exits non-zero should be recreated with an increasing delay (e.g. the same 5s..80s ladder as failedRunnerBackoff, or capped at a few minutes), resetting once a pod reaches Running. A fleet-wide auth failure then degrades into a slow retry instead of a request storm, and kubectl get pods shows a stable Error pod long enough to inspect.
Proposed fix (happy to open a PR with tests if the direction is acceptable):
- Mirror
EphemeralRunnerStatus: add Failures map[string]metav1.Time (or a count + last-failure time) to AutoscalingListenerStatus (currently an empty struct).
- In the
cs.State.Terminated != nil branch, when ExitCode != 0, record the failure, compute delay := listenerRestartBackoff(len(failures), cs.State.Terminated.FinishedAt.Time, now), and if delay > 0 return ctrl.Result{RequeueAfter: delay} before calling deleteListenerPod, leaving the terminated pod in place as the visible backoff marker. When the delay has elapsed, delete as today.
- In the
cs.State.Running != nil branch, clear the failures.
- A pure
listenerRestartBackoff(failures int, finishedAt, now time.Time) time.Duration helper keeps this unit-testable without envtest; one extra envtest case can patch the pod status to Terminated{ExitCode: 1, FinishedAt: now} and assert the pod is not deleted inside the window, then is.
A minimal, CRD-free variant is also possible: enforce a fixed minimum delay (e.g. 10-15 s) measured from Terminated.FinishedAt before deleting. It caps the storm without a counter, at the cost of no escalation.
Additional Context
# Nothing unusual is required. Any gha-runner-scale-set values with a githubConfigSecret whose token is invalid reproduce it, e.g.:
githubConfigUrl: https://github.com/<org>
githubConfigSecret: <secret-with-revoked-pat>
minRunners: 0
maxRunners: 5
Controller Logs
# Sequence emitted for one scale set, repeating indefinitely with sub-second spacing between the delete and the next create.
# (Log messages as emitted by autoscalinglistener_controller.go at the linked lines; a full controller-manager log gist can be provided on request.)
INFO AutoscalingListener Listener pod is terminated {"namespace": "<ns>", "name": "<scaleset>-<hash>-listener", "reason": "Error", "message": ""}
INFO AutoscalingListener Deleting the listener pod {"namespace": "<ns>", "name": "<scaleset>-<hash>-listener"}
INFO AutoscalingListener Creating listener pod {"namespace": "<ns>", "name": "<scaleset>-<hash>-listener"}
INFO AutoscalingListener Listener pod is not ready {"namespace": "<ns>", "name": "<scaleset>-<hash>-listener"}
INFO AutoscalingListener Listener pod is terminated ...
Runner Pod Logs
# Listener pod (the only pod involved; no runner pods are created while the listener cannot start). Typical last line before exit 1:
INFO listener-app getting runner registration token {"registrationTokenURL": "https://api.github.com/orgs/<org>/actions/runners/registration-token"}
Application returned an error: failed to create actions client: ... StatusCode 401 ... Bad credentials
# or, once REST quota is gone:
Application returned an error: ... StatusCode 403 ... API rate limit exceeded for user ID <id>
Analysis prepared with an AI agent operated by KR-Ravindra, who verified the code paths.
Checks
Controller Version
0.14.2 (same code path on
masterat 03328aa)Deployment Method
Helm
Checks
To Reproduce
Describe the bug
The
AutoscalingListenerreconciler has no failure accounting and no delay between a listener pod terminating and the next pod being created:Terminated->deleteListenerPodimmediately,ctrl.Result{}:actions-runner-controller/controllers/actions.github.com/autoscalinglistener_controller.go
Lines 547 to 556 in 03328aa
NotFound->r.Create(desiredPod)immediately:actions-runner-controller/controllers/actions.github.com/autoscalinglistener_controller.go
Lines 502 to 528 in 03328aa
deleteListenerPod:actions-runner-controller/controllers/actions.github.com/autoscalinglistener_controller.go
Lines 572 to 585 in 03328aa
The listener pod is built with
RestartPolicy: Never(actions-runner-controller/controllers/actions.github.com/resourcebuilder.go
Line 437 in 03328aa
Owns(&corev1.Pod{})watch events, which enqueue with a plainAddrather thanAddRateLimited, so the workqueue's per-item exponential backoff does not apply (it only applies to requeues/errors). Neither--workqueue-rate-limitersetting changes this.Contrast with the
EphemeralRunnerreconciler, which records failures in status and backs off 5s -> 10s -> 20s -> 40s -> 80s before recreating a runner pod (#4059):actions-runner-controller/controllers/actions.github.com/ephemeralrunner_controller.go
Lines 65 to 76 in 03328aa
actions-runner-controller/controllers/actions.github.com/ephemeralrunner_controller.go
Lines 258 to 274 in 03328aa
The listener gets no equivalent. The existing test
It should re-create pod but persist config secret whenever listener container is terminated(actions-runner-controller/controllers/actions.github.com/autoscalinglistener_controller_test.go
Line 576 in 03328aa
Why it matters: every listener start repeats the full auth handshake against GitHub (installation/registration token, then the Actions admin connection; the admin-connection client itself retries 401/403 up to 4 more times in-process) before exiting. When the failure is deterministic (expired/rotated PAT, exhausted REST rate limit, scale set gone server-side), every listener in the cluster enters a create/crash/delete loop of a few seconds and the loop itself burns the remaining quota. In a cluster with several hundred scale sets on one PAT we saw several hundred listener pods in
Error, sustained 403API rate limit exceeded, andkubectl get podsoften showed no listener pod at all because each one lived 1-3 s. Recovery after fixing the credential was delayed because the loop had kept the limit pinned. #3942 asked for in-process retries and was closed with "the listener would try to restart and recover on its own" -- that restart path is the intended recovery mechanism, so it should be paced.Describe the expected behavior
A listener pod that exits non-zero should be recreated with an increasing delay (e.g. the same 5s..80s ladder as
failedRunnerBackoff, or capped at a few minutes), resetting once a pod reachesRunning. A fleet-wide auth failure then degrades into a slow retry instead of a request storm, andkubectl get podsshows a stableErrorpod long enough to inspect.Proposed fix (happy to open a PR with tests if the direction is acceptable):
EphemeralRunnerStatus: addFailures map[string]metav1.Time(or a count + last-failure time) toAutoscalingListenerStatus(currently an empty struct).cs.State.Terminated != nilbranch, whenExitCode != 0, record the failure, computedelay := listenerRestartBackoff(len(failures), cs.State.Terminated.FinishedAt.Time, now), and ifdelay > 0returnctrl.Result{RequeueAfter: delay}before callingdeleteListenerPod, leaving the terminated pod in place as the visible backoff marker. When the delay has elapsed, delete as today.cs.State.Running != nilbranch, clear the failures.listenerRestartBackoff(failures int, finishedAt, now time.Time) time.Durationhelper keeps this unit-testable without envtest; one extra envtest case can patch the pod status toTerminated{ExitCode: 1, FinishedAt: now}and assert the pod is not deleted inside the window, then is.A minimal, CRD-free variant is also possible: enforce a fixed minimum delay (e.g. 10-15 s) measured from
Terminated.FinishedAtbefore deleting. It caps the storm without a counter, at the cost of no escalation.Additional Context
Controller Logs
Runner Pod Logs
Analysis prepared with an AI agent operated by KR-Ravindra, who verified the code paths.