Skip to content

Support Configurable Max Exponential Backoff Delay for Direct Controllers #13649

Description

@maqiuyujoyce

Problem Statement

Under default controller-runtime reconciliation loops, any reconciliation failure causes the resource to enter an exponential backoff retry loop. In Config Connector, this retry backoff is capped by default at 120 seconds (pkg/controller/ratelimiter/ratelimiter.go).

For GCP services with strict API rate limits or resources encountering prolonged non-transient failures, retrying every ~2 minutes indefinitely burns GCP project API quota, spams Cloud Audit Logs, and causes unnecessary workqueue contention. Currently, operators have no declarative way to relax error retry pacing or halt automated retries on an individual failing direct resource without affecting sibling resources.


Proposed Solution

Introduce the resource-level annotation:
cnrm.cloud.google.com/backoff-max-delay-in-seconds

This annotation configures the maximum delay ceiling for exponential backoff during error retries on the target direct resource.

Specification & Semantics

  1. Format: Non-negative string integer representing the maximum delay ceiling in seconds (e.g., "600" for 10m, "1800" for 30m).
  2. Default Ceiling: If the annotation is not set, Config Connector retains the current default ceiling (120s).
  3. Special Value "0" (Complete Error Halt):
    • Setting backoff-max-delay-in-seconds: "0" instructs Config Connector to completely stop automated re-reconciliation for this resource upon error.
    • After updating .status.conditions[Ready]=False with the failure details, the reconciler returns reconcile.Result{}, nil.
    • No periodic re-enqueue: Because reconcile.Result{}, nil is returned (and no RequeueAfter is scheduled), controller-runtime will not re-enqueue the object after the standard 10–20 minute drift period. It stays completely halted until:
      • The user updates .spec (triggering an immediate watch event via Queue.Add(req)).
      • The user modifies annotations or metadata.
      • The Config Connector controller pod restarts.
    • This provides parity with cnrm.cloud.google.com/reconcile-interval-in-seconds: "0" for resources in an error state.

Low-Level Design Details

1. Constants (pkg/k8s/constants.go)

Declare the annotation key and base default values:

const (
    BackoffMaxDelayInSecondsAnnotation = "cnrm.cloud.google.com/backoff-max-delay-in-seconds"
    DefaultBackoffMaxDelay             = 120 * time.Second
    DefaultBackoffBaseDelay            = 1 * time.Second
)

2. Dedicated Dynamic Rate Limiter & Parsing (pkg/controller/ratelimiter/dynamicratelimiter.go)

Create pkg/controller/ratelimiter/dynamicratelimiter.go to encapsulate parsing, validation, and per-resource failure tracking and backoff calculations:

package ratelimiter

import (
    "fmt"
    "math"
    "strconv"
    "sync"
    "time"

    "github.com/GoogleCloudPlatform/k8s-config-connector/pkg/k8s"
    "k8s.io/apimachinery/pkg/types"
)

// ParseBackoffMaxDelay extracts and validates the annotation value.
// Returns:
//   - (nil, nil) if annotation is unset or empty.
//   - (ptr(0), nil) if set to "0" (halt retries).
//   - (ptr(duration), nil) if set to a positive integer.
//   - (nil, error) if value is negative or not a valid integer.
func ParseBackoffMaxDelay(annotations map[string]string) (*time.Duration, error) {
    val, ok := annotations[k8s.BackoffMaxDelayInSecondsAnnotation]
    if !ok || val == "" {
        return nil, nil
    }
    seconds, err := strconv.ParseInt(val, 10, 64)
    if err != nil {
        return nil, fmt.Errorf("invalid value %q for annotation %s: must be an integer", val, k8s.BackoffMaxDelayInSecondsAnnotation)
    }
    if seconds < 0 {
        return nil, fmt.Errorf("invalid value %q for annotation %s: must be non-negative", val, k8s.BackoffMaxDelayInSecondsAnnotation)
    }
    d := time.Duration(seconds) * time.Second
    return &d, nil
}

// DynamicRateLimiter tracks consecutive reconciliation failures per resource
// and calculates exponential retry intervals clamped to the resource's max delay.
type DynamicRateLimiter struct {
    mu       sync.Mutex
    failures map[types.NamespacedName]int
}

func NewDynamicRateLimiter() *DynamicRateLimiter {
    return &DynamicRateLimiter{
        failures: make(map[types.NamespacedName]int),
    }
}

// NextDelay increments the failure count and calculates the exponential backoff:
// baseDelay * 2^(failures-1), capped at maxDelay.
func (r *DynamicRateLimiter) NextDelay(nn types.NamespacedName, baseDelay, maxDelay time.Duration) time.Duration {
    r.mu.Lock()
    defer r.mu.Unlock()

    exp := r.failures[nn]
    r.failures[nn]++

    backoff := float64(baseDelay) * math.Pow(2, float64(exp))
    if backoff > float64(maxDelay) || backoff <= 0 {
        return maxDelay
    }
    return time.Duration(backoff)
}

// Forget clears the failure count when reconciliation succeeds (matching controller-runtime semantics).
func (r *DynamicRateLimiter) Forget(nn types.NamespacedName) {
    r.mu.Lock()
    defer r.mu.Unlock()
    delete(r.failures, nn)
}

3. Direct Controller Isolation (pkg/controller/direct/directbase/directbase_controller.go)

Because ParentReconciler (pkg/controller/parent/controller.go) manages all underlying controller types, we do not modify parent/controller.go or the parent workqueue rate limiter.

Instead, DirectReconciler embeds DynamicRateLimiter and controls its own return values. ParentReconciler transparently passes DirectReconciler's reconcile.Result back to controller-runtime:

In DirectReconciler struct:

type DirectReconciler struct {
    ...
    rateLimiter *ratelimiter.DynamicRateLimiter
}

In NewReconciler(...):

r := DirectReconciler{
    ...
    rateLimiter: ratelimiter.NewDynamicRateLimiter(),
}

In Reconcile(...):

requeue, err := runCtx.doReconcile(ctx, obj)
if err != nil {
    // 1. Deletion error: do NOT apply custom ceiling or "0" halt.
    // Return err so controller-runtime uses standard workqueue rate limiting to retry deletion.
    if obj.GetDeletionTimestamp() != nil {
        return reconcile.Result{}, err
    }

    maxDelay, parseErr := ratelimiter.ParseBackoffMaxDelay(obj.GetAnnotations())
    if parseErr != nil {
        logger.Error(parseErr, "invalid backoff-max-delay annotation", "resource", request.NamespacedName)
        // Fall back to default rate limiting behavior
        return reconcile.Result{}, err
    }

    // 2. Special value "0": Halt automated error retries
    if maxDelay != nil && *maxDelay == 0 {
        logger.Info("Halting automated error retries as backoff-max-delay-in-seconds is set to 0",
            "resource", request.NamespacedName,
            "error", err)
        // Returning Result{}, nil stops workqueue requeueing and does not schedule jitteredPeriod
        return reconcile.Result{}, nil
    }

    // 3. Custom max delay ceiling: calculate exponential backoff up to maxDelay
    if maxDelay != nil && *maxDelay > 0 {
        nextDelay := r.rateLimiter.NextDelay(
            request.NamespacedName,
            k8s.DefaultBackoffBaseDelay,
            *maxDelay,
        )
        logger.Info("Scheduling error retry with custom max delay ceiling",
            "resource", request.NamespacedName,
            "retryAfter", nextDelay,
            "maxDelay", *maxDelay)
        return reconcile.Result{RequeueAfter: nextDelay}, nil
    }

    // 4. Default: delegate to controller-runtime workqueue rate limiter (120s ceiling)
    return reconcile.Result{}, err
}

// Reconcile succeeded: clear failure counter (standard controller-runtime Forget)
r.rateLimiter.Forget(request.NamespacedName)

4. Workqueue & Deletion Lifecycle Guarantees

  1. Immediate Wakeup on .spec Changes: KCC's UnderlyingResourceOutOfSyncPredicate detects .spec generation increments and enqueues the resource via Queue.Add(req). In client-go, calling Queue.Add(req) immediately promotes the item into the active ready queue, bypassing any pending delay timer.
  2. Standard Reset on Success: Like standard controller-runtime, r.rateLimiter.Forget(request.NamespacedName) is called on every successful reconciliation, wiping the failure counter to 0.
  3. Deletion Intent Precedence: When obj.GetDeletionTimestamp() != nil, all backoff calculations are bypassed to guarantee that kubectl delete teardown reconciliations are never delayed.
  4. Zero Impact on TF/DCL: Legacy controllers continue returning standard Result{}, err and use the default workqueue limiter without disruption.

E2E Scenario Test Plan

Add an end-to-end scenario test in tests/e2e/testdata/scenarios/backoff_max_delay_zero/script.yaml using a lightweight direct resource (such as LoggingLogMetric). Use record-real-gcp skill to record real GCP logs:

  1. Step 0 — Create Failing Resource with backoff-max-delay-in-seconds: "0":
    Apply a LoggingLogMetric with an invalid filter query and backoff-max-delay-in-seconds: "0".
    • Verification: Initial reconcile fails against GCP, and status updates to Ready=False (reason: UpdateFailed).
  2. Step 1 — Verify Zero Additional Retries (TEST: APPLY-NO-WAIT):
    Wait 10–15 seconds without updating the resource.
    • Verification: Zero subsequent HTTP calls are made to GCP, confirming reconcile.Result{}, nil prevented workqueue re-enqueuing.
  3. Step 2 — Fix .spec & Immediate Recovery (TEST: WAIT-FOR-HTTP-REQUEST):
    Update .spec with a valid filter expression.
    • Verification: Generation change triggers immediate reconciliation against GCP and status transitions to Ready=True.
  4. Step 3 — Teardown (TEST: DELETE):
    Delete the resource and verify finalizer removal and clean deletion.

Acceptance Criteria

  • pkg/controller/ratelimiter/dynamicratelimiter.go implemented with ParseBackoffMaxDelay and DynamicRateLimiter.
  • Setting backoff-max-delay-in-seconds: "0" logs the halt, records Ready=False in .status.conditions, and returns reconcile.Result{}, nil.
  • Setting backoff-max-delay-in-seconds: "0" does not re-enqueue after 10–20 minutes automatically.
  • Setting backoff-max-delay-in-seconds: "<seconds>" (e.g., "600") caps exponential retry delay at the specified ceiling.
  • Updating .spec immediately interrupts any pending backoff delay and triggers reconciliation.
  • Resources in deletion (deletionTimestamp != nil) bypass error retry delays.
  • Comprehensive unit tests in pkg/controller/ratelimiter/dynamicratelimiter_test.go and pkg/controller/direct/directbase/directbase_controller_test.go covering "0", positive integers, malformed input, success resets, and deletion bypass.
  • E2E scenario test in tests/e2e/testdata/scenarios/backoff_max_delay_zero/ passing against both real GCP and MockGCP.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions