Skip to content

AutoscalingListener recreates crashed listener pod immediately with no backoff #4656

Description

@KR-Ravindra

Checks

Controller Version

0.14.2 (same code path on master at 03328aa)

Deployment Method

Helm

Checks

  • This isn't a question or user support case (For Q&A and community support, go to Discussions).
  • I've read the Changelog before submitting this issue and I'm sure it's not due to any recently-introduced backward-incompatible changes

To Reproduce

1. Install gha-runner-scale-set-controller and one gha-runner-scale-set with a `githubConfigSecret` whose token is revoked or expired (any deterministic startup failure works: 401 from the registration-token API, 403 from REST rate-limit exhaustion, 404 RunnerScaleSetNotFoundException).
2. Watch the listener pod in the controller namespace: `kubectl get pods -n <controller-ns> -w`.
3. The pod is created, runs the auth handshake, exits non-zero within a few seconds, is deleted by the controller and recreated immediately. It never reaches a steady failed state and the loop never slows down.
4. With N scale sets installed, all N listeners loop in lock-step. Fixing the secret ends the loop, but while it runs, each iteration spends more GitHub API quota.

Describe the bug

The AutoscalingListener reconciler has no failure accounting and no delay between a listener pod terminating and the next pod being created:

  • container Terminated -> deleteListenerPod immediately, ctrl.Result{}:
    case cs.State.Terminated != nil:
    log.Info(
    "Listener pod is terminated",
    "namespace", listenerPod.Namespace,
    "name", listenerPod.Name,
    "reason", cs.State.Terminated.Reason,
    "message", cs.State.Terminated.Message,
    )
    return ctrl.Result{}, r.deleteListenerPod(ctx, &autoscalingListener, &listenerPod, log)
  • pod NotFound -> r.Create(desiredPod) immediately:
    case kerrors.IsNotFound(err):
    if err := r.publishRunningListener(&autoscalingListener, false); err != nil {
    // If publish fails, URL is incorrect which means the listener pod would never be able to start
    return ctrl.Result{}, nil
    }
    r.ResourceCache.listenerPod.Delete(&autoscalingListener)
    desiredPod, err := r.newScaleSetListenerPod(
    &autoscalingListener,
    &listenerConfigSecret,
    &serviceAccount,
    &listenerRole,
    &listenerRoleBinding,
    metricsConfig,
    )
    if err != nil {
    log.Error(err, "Failed to build listener pod")
    return ctrl.Result{}, err
    }
    log.Info("Creating listener pod", "namespace", desiredPod.Namespace, "name", desiredPod.Name)
    if err := r.Create(ctx, desiredPod); err != nil {
    log.Error(err, "Unable to create listener pod", "namespace", desiredPod.Namespace, "name", desiredPod.Name)
    return ctrl.Result{}, err
    }
    default: // error
    log.Error(err, "Unable to get listener pod", "namespace", autoscalingListener.Namespace, "name", autoscalingListener.Name)
  • deleteListenerPod:
    func (r *AutoscalingListenerReconciler) deleteListenerPod(ctx context.Context, autoscalingListener *v1alpha1.AutoscalingListener, listenerPod *corev1.Pod, log logr.Logger) error {
    if err := r.publishRunningListener(autoscalingListener, false); err != nil {
    log.Error(err, "Unable to publish runner listener down metric", "namespace", listenerPod.Namespace, "name", listenerPod.Name)
    }
    if listenerPod.DeletionTimestamp.IsZero() {
    log.Info("Deleting the listener pod", "namespace", listenerPod.Namespace, "name", listenerPod.Name)
    if err := r.Delete(ctx, listenerPod); err != nil && !kerrors.IsNotFound(err) {
    log.Error(err, "Unable to delete the listener pod", "namespace", listenerPod.Namespace, "name", listenerPod.Name)
    return err
    }
    }
    return nil
    }

The listener pod is built with RestartPolicy: Never (

RestartPolicy: corev1.RestartPolicyNever,
), so the kubelet's CrashLoopBackOff never applies either; the controller is the only thing pacing restarts. Both transitions are driven by Owns(&corev1.Pod{}) watch events, which enqueue with a plain Add rather than AddRateLimited, so the workqueue's per-item exponential backoff does not apply (it only applies to requeues/errors). Neither --workqueue-rate-limiter setting changes this.

Contrast with the EphemeralRunner reconciler, which records failures in status and backs off 5s -> 10s -> 20s -> 40s -> 80s before recreating a runner pod (#4059):

  • // precompute backoff durations for failed ephemeral runners
    // the len(failedRunnerBackoff) must be equal to maxFailures + 1
    var failedRunnerBackoff = []time.Duration{
    0,
    5 * time.Second,
    10 * time.Second,
    20 * time.Second,
    40 * time.Second,
    80 * time.Second,
    }
    const maxFailures = 5
  • now := metav1.Now()
    lastFailure := ephemeralRunner.Status.LastFailure()
    backoffDuration := failedRunnerBackoff[len(ephemeralRunner.Status.Failures)]
    nextReconciliation := lastFailure.Add(backoffDuration)
    if !lastFailure.IsZero() && now.Before(&metav1.Time{Time: nextReconciliation}) {
    requeueAfter := nextReconciliation.Sub(now.Time)
    log.Info(
    "Backing off the next reconciliation due to failure",
    "lastFailure", lastFailure,
    "nextReconciliation", nextReconciliation,
    "requeueAfter", requeueAfter,
    )
    return ctrl.Result{
    Requeue: true,
    RequeueAfter: requeueAfter,
    }, nil
    }

The listener gets no equivalent. The existing test It should re-create pod but persist config secret whenever listener container is terminated (

It("It should re-create pod but persist config secret whenever listener container is terminated", func() {
) pins the current behaviour: recreation is immediate and unconditional, even for exit code 0.

Why it matters: every listener start repeats the full auth handshake against GitHub (installation/registration token, then the Actions admin connection; the admin-connection client itself retries 401/403 up to 4 more times in-process) before exiting. When the failure is deterministic (expired/rotated PAT, exhausted REST rate limit, scale set gone server-side), every listener in the cluster enters a create/crash/delete loop of a few seconds and the loop itself burns the remaining quota. In a cluster with several hundred scale sets on one PAT we saw several hundred listener pods in Error, sustained 403 API rate limit exceeded, and kubectl get pods often showed no listener pod at all because each one lived 1-3 s. Recovery after fixing the credential was delayed because the loop had kept the limit pinned. #3942 asked for in-process retries and was closed with "the listener would try to restart and recover on its own" -- that restart path is the intended recovery mechanism, so it should be paced.

Describe the expected behavior

A listener pod that exits non-zero should be recreated with an increasing delay (e.g. the same 5s..80s ladder as failedRunnerBackoff, or capped at a few minutes), resetting once a pod reaches Running. A fleet-wide auth failure then degrades into a slow retry instead of a request storm, and kubectl get pods shows a stable Error pod long enough to inspect.

Proposed fix (happy to open a PR with tests if the direction is acceptable):

  1. Mirror EphemeralRunnerStatus: add Failures map[string]metav1.Time (or a count + last-failure time) to AutoscalingListenerStatus (currently an empty struct).
  2. In the cs.State.Terminated != nil branch, when ExitCode != 0, record the failure, compute delay := listenerRestartBackoff(len(failures), cs.State.Terminated.FinishedAt.Time, now), and if delay > 0 return ctrl.Result{RequeueAfter: delay} before calling deleteListenerPod, leaving the terminated pod in place as the visible backoff marker. When the delay has elapsed, delete as today.
  3. In the cs.State.Running != nil branch, clear the failures.
  4. A pure listenerRestartBackoff(failures int, finishedAt, now time.Time) time.Duration helper keeps this unit-testable without envtest; one extra envtest case can patch the pod status to Terminated{ExitCode: 1, FinishedAt: now} and assert the pod is not deleted inside the window, then is.

A minimal, CRD-free variant is also possible: enforce a fixed minimum delay (e.g. 10-15 s) measured from Terminated.FinishedAt before deleting. It caps the storm without a counter, at the cost of no escalation.

Additional Context

# Nothing unusual is required. Any gha-runner-scale-set values with a githubConfigSecret whose token is invalid reproduce it, e.g.:
githubConfigUrl: https://github.com/<org>
githubConfigSecret: <secret-with-revoked-pat>
minRunners: 0
maxRunners: 5

Controller Logs

# Sequence emitted for one scale set, repeating indefinitely with sub-second spacing between the delete and the next create.
# (Log messages as emitted by autoscalinglistener_controller.go at the linked lines; a full controller-manager log gist can be provided on request.)
INFO  AutoscalingListener  Listener pod is terminated   {"namespace": "<ns>", "name": "<scaleset>-<hash>-listener", "reason": "Error", "message": ""}
INFO  AutoscalingListener  Deleting the listener pod    {"namespace": "<ns>", "name": "<scaleset>-<hash>-listener"}
INFO  AutoscalingListener  Creating listener pod        {"namespace": "<ns>", "name": "<scaleset>-<hash>-listener"}
INFO  AutoscalingListener  Listener pod is not ready    {"namespace": "<ns>", "name": "<scaleset>-<hash>-listener"}
INFO  AutoscalingListener  Listener pod is terminated   ...

Runner Pod Logs

# Listener pod (the only pod involved; no runner pods are created while the listener cannot start). Typical last line before exit 1:
INFO  listener-app  getting runner registration token  {"registrationTokenURL": "https://api.github.com/orgs/<org>/actions/runners/registration-token"}
Application returned an error: failed to create actions client: ... StatusCode 401 ... Bad credentials
# or, once REST quota is gone:
Application returned an error: ... StatusCode 403 ... API rate limit exceeded for user ID <id>

Analysis prepared with an AI agent operated by KR-Ravindra, who verified the code paths.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions