Problem Statement
Under default controller-runtime reconciliation loops, any reconciliation failure causes the resource to enter an exponential backoff retry loop. In Config Connector, this retry backoff is capped by default at 120 seconds (pkg/controller/ratelimiter/ratelimiter.go).
For GCP services with strict API rate limits or resources encountering prolonged non-transient failures, retrying every ~2 minutes indefinitely burns GCP project API quota, spams Cloud Audit Logs, and causes unnecessary workqueue contention. Currently, operators have no declarative way to relax error retry pacing or halt automated retries on an individual failing direct resource without affecting sibling resources.
Proposed Solution
Introduce the resource-level annotation:
cnrm.cloud.google.com/backoff-max-delay-in-seconds
This annotation configures the maximum delay ceiling for exponential backoff during error retries on the target direct resource.
Specification & Semantics
- Format: Non-negative string integer representing the maximum delay ceiling in seconds (e.g.,
"600" for 10m, "1800" for 30m).
- Default Ceiling: If the annotation is not set, Config Connector retains the current default ceiling (120s).
- Special Value
"0" (Complete Error Halt):
- Setting
backoff-max-delay-in-seconds: "0" instructs Config Connector to completely stop automated re-reconciliation for this resource upon error.
- After updating
.status.conditions[Ready]=False with the failure details, the reconciler returns reconcile.Result{}, nil.
- No periodic re-enqueue: Because
reconcile.Result{}, nil is returned (and no RequeueAfter is scheduled), controller-runtime will not re-enqueue the object after the standard 10–20 minute drift period. It stays completely halted until:
- The user updates
.spec (triggering an immediate watch event via Queue.Add(req)).
- The user modifies annotations or metadata.
- The Config Connector controller pod restarts.
- This provides parity with
cnrm.cloud.google.com/reconcile-interval-in-seconds: "0" for resources in an error state.
Low-Level Design Details
1. Constants (pkg/k8s/constants.go)
Declare the annotation key and base default values:
const (
BackoffMaxDelayInSecondsAnnotation = "cnrm.cloud.google.com/backoff-max-delay-in-seconds"
DefaultBackoffMaxDelay = 120 * time.Second
DefaultBackoffBaseDelay = 1 * time.Second
)
2. Dedicated Dynamic Rate Limiter & Parsing (pkg/controller/ratelimiter/dynamicratelimiter.go)
Create pkg/controller/ratelimiter/dynamicratelimiter.go to encapsulate parsing, validation, and per-resource failure tracking and backoff calculations:
package ratelimiter
import (
"fmt"
"math"
"strconv"
"sync"
"time"
"github.com/GoogleCloudPlatform/k8s-config-connector/pkg/k8s"
"k8s.io/apimachinery/pkg/types"
)
// ParseBackoffMaxDelay extracts and validates the annotation value.
// Returns:
// - (nil, nil) if annotation is unset or empty.
// - (ptr(0), nil) if set to "0" (halt retries).
// - (ptr(duration), nil) if set to a positive integer.
// - (nil, error) if value is negative or not a valid integer.
func ParseBackoffMaxDelay(annotations map[string]string) (*time.Duration, error) {
val, ok := annotations[k8s.BackoffMaxDelayInSecondsAnnotation]
if !ok || val == "" {
return nil, nil
}
seconds, err := strconv.ParseInt(val, 10, 64)
if err != nil {
return nil, fmt.Errorf("invalid value %q for annotation %s: must be an integer", val, k8s.BackoffMaxDelayInSecondsAnnotation)
}
if seconds < 0 {
return nil, fmt.Errorf("invalid value %q for annotation %s: must be non-negative", val, k8s.BackoffMaxDelayInSecondsAnnotation)
}
d := time.Duration(seconds) * time.Second
return &d, nil
}
// DynamicRateLimiter tracks consecutive reconciliation failures per resource
// and calculates exponential retry intervals clamped to the resource's max delay.
type DynamicRateLimiter struct {
mu sync.Mutex
failures map[types.NamespacedName]int
}
func NewDynamicRateLimiter() *DynamicRateLimiter {
return &DynamicRateLimiter{
failures: make(map[types.NamespacedName]int),
}
}
// NextDelay increments the failure count and calculates the exponential backoff:
// baseDelay * 2^(failures-1), capped at maxDelay.
func (r *DynamicRateLimiter) NextDelay(nn types.NamespacedName, baseDelay, maxDelay time.Duration) time.Duration {
r.mu.Lock()
defer r.mu.Unlock()
exp := r.failures[nn]
r.failures[nn]++
backoff := float64(baseDelay) * math.Pow(2, float64(exp))
if backoff > float64(maxDelay) || backoff <= 0 {
return maxDelay
}
return time.Duration(backoff)
}
// Forget clears the failure count when reconciliation succeeds (matching controller-runtime semantics).
func (r *DynamicRateLimiter) Forget(nn types.NamespacedName) {
r.mu.Lock()
defer r.mu.Unlock()
delete(r.failures, nn)
}
3. Direct Controller Isolation (pkg/controller/direct/directbase/directbase_controller.go)
Because ParentReconciler (pkg/controller/parent/controller.go) manages all underlying controller types, we do not modify parent/controller.go or the parent workqueue rate limiter.
Instead, DirectReconciler embeds DynamicRateLimiter and controls its own return values. ParentReconciler transparently passes DirectReconciler's reconcile.Result back to controller-runtime:
In DirectReconciler struct:
type DirectReconciler struct {
...
rateLimiter *ratelimiter.DynamicRateLimiter
}
In NewReconciler(...):
r := DirectReconciler{
...
rateLimiter: ratelimiter.NewDynamicRateLimiter(),
}
In Reconcile(...):
requeue, err := runCtx.doReconcile(ctx, obj)
if err != nil {
// 1. Deletion error: do NOT apply custom ceiling or "0" halt.
// Return err so controller-runtime uses standard workqueue rate limiting to retry deletion.
if obj.GetDeletionTimestamp() != nil {
return reconcile.Result{}, err
}
maxDelay, parseErr := ratelimiter.ParseBackoffMaxDelay(obj.GetAnnotations())
if parseErr != nil {
logger.Error(parseErr, "invalid backoff-max-delay annotation", "resource", request.NamespacedName)
// Fall back to default rate limiting behavior
return reconcile.Result{}, err
}
// 2. Special value "0": Halt automated error retries
if maxDelay != nil && *maxDelay == 0 {
logger.Info("Halting automated error retries as backoff-max-delay-in-seconds is set to 0",
"resource", request.NamespacedName,
"error", err)
// Returning Result{}, nil stops workqueue requeueing and does not schedule jitteredPeriod
return reconcile.Result{}, nil
}
// 3. Custom max delay ceiling: calculate exponential backoff up to maxDelay
if maxDelay != nil && *maxDelay > 0 {
nextDelay := r.rateLimiter.NextDelay(
request.NamespacedName,
k8s.DefaultBackoffBaseDelay,
*maxDelay,
)
logger.Info("Scheduling error retry with custom max delay ceiling",
"resource", request.NamespacedName,
"retryAfter", nextDelay,
"maxDelay", *maxDelay)
return reconcile.Result{RequeueAfter: nextDelay}, nil
}
// 4. Default: delegate to controller-runtime workqueue rate limiter (120s ceiling)
return reconcile.Result{}, err
}
// Reconcile succeeded: clear failure counter (standard controller-runtime Forget)
r.rateLimiter.Forget(request.NamespacedName)
4. Workqueue & Deletion Lifecycle Guarantees
- Immediate Wakeup on
.spec Changes: KCC's UnderlyingResourceOutOfSyncPredicate detects .spec generation increments and enqueues the resource via Queue.Add(req). In client-go, calling Queue.Add(req) immediately promotes the item into the active ready queue, bypassing any pending delay timer.
- Standard Reset on Success: Like standard
controller-runtime, r.rateLimiter.Forget(request.NamespacedName) is called on every successful reconciliation, wiping the failure counter to 0.
- Deletion Intent Precedence: When
obj.GetDeletionTimestamp() != nil, all backoff calculations are bypassed to guarantee that kubectl delete teardown reconciliations are never delayed.
- Zero Impact on TF/DCL: Legacy controllers continue returning standard
Result{}, err and use the default workqueue limiter without disruption.
E2E Scenario Test Plan
Add an end-to-end scenario test in tests/e2e/testdata/scenarios/backoff_max_delay_zero/script.yaml using a lightweight direct resource (such as LoggingLogMetric). Use record-real-gcp skill to record real GCP logs:
- Step 0 — Create Failing Resource with
backoff-max-delay-in-seconds: "0":
Apply a LoggingLogMetric with an invalid filter query and backoff-max-delay-in-seconds: "0".
- Verification: Initial reconcile fails against GCP, and status updates to
Ready=False (reason: UpdateFailed).
- Step 1 — Verify Zero Additional Retries (
TEST: APPLY-NO-WAIT):
Wait 10–15 seconds without updating the resource.
- Verification: Zero subsequent HTTP calls are made to GCP, confirming
reconcile.Result{}, nil prevented workqueue re-enqueuing.
- Step 2 — Fix
.spec & Immediate Recovery (TEST: WAIT-FOR-HTTP-REQUEST):
Update .spec with a valid filter expression.
- Verification: Generation change triggers immediate reconciliation against GCP and status transitions to
Ready=True.
- Step 3 — Teardown (
TEST: DELETE):
Delete the resource and verify finalizer removal and clean deletion.
Acceptance Criteria
Problem Statement
Under default
controller-runtimereconciliation loops, any reconciliation failure causes the resource to enter an exponential backoff retry loop. In Config Connector, this retry backoff is capped by default at 120 seconds (pkg/controller/ratelimiter/ratelimiter.go).For GCP services with strict API rate limits or resources encountering prolonged non-transient failures, retrying every ~2 minutes indefinitely burns GCP project API quota, spams Cloud Audit Logs, and causes unnecessary workqueue contention. Currently, operators have no declarative way to relax error retry pacing or halt automated retries on an individual failing direct resource without affecting sibling resources.
Proposed Solution
Introduce the resource-level annotation:
cnrm.cloud.google.com/backoff-max-delay-in-secondsThis annotation configures the maximum delay ceiling for exponential backoff during error retries on the target direct resource.
Specification & Semantics
"600"for 10m,"1800"for 30m)."0"(Complete Error Halt):backoff-max-delay-in-seconds: "0"instructs Config Connector to completely stop automated re-reconciliation for this resource upon error..status.conditions[Ready]=Falsewith the failure details, the reconciler returnsreconcile.Result{}, nil.reconcile.Result{}, nilis returned (and noRequeueAfteris scheduled),controller-runtimewill not re-enqueue the object after the standard 10–20 minute drift period. It stays completely halted until:.spec(triggering an immediate watch event viaQueue.Add(req)).cnrm.cloud.google.com/reconcile-interval-in-seconds: "0"for resources in an error state.Low-Level Design Details
1. Constants (
pkg/k8s/constants.go)Declare the annotation key and base default values:
2. Dedicated Dynamic Rate Limiter & Parsing (
pkg/controller/ratelimiter/dynamicratelimiter.go)Create
pkg/controller/ratelimiter/dynamicratelimiter.goto encapsulate parsing, validation, and per-resource failure tracking and backoff calculations:3. Direct Controller Isolation (
pkg/controller/direct/directbase/directbase_controller.go)Because
ParentReconciler(pkg/controller/parent/controller.go) manages all underlying controller types, we do not modifyparent/controller.goor the parent workqueue rate limiter.Instead,
DirectReconcilerembedsDynamicRateLimiterand controls its own return values.ParentReconcilertransparently passesDirectReconciler'sreconcile.Resultback tocontroller-runtime:In
DirectReconcilerstruct:In
NewReconciler(...):In
Reconcile(...):4. Workqueue & Deletion Lifecycle Guarantees
.specChanges: KCC'sUnderlyingResourceOutOfSyncPredicatedetects.specgeneration increments and enqueues the resource viaQueue.Add(req). Inclient-go, callingQueue.Add(req)immediately promotes the item into the active ready queue, bypassing any pending delay timer.controller-runtime,r.rateLimiter.Forget(request.NamespacedName)is called on every successful reconciliation, wiping the failure counter to 0.obj.GetDeletionTimestamp() != nil, all backoff calculations are bypassed to guarantee thatkubectl deleteteardown reconciliations are never delayed.Result{}, errand use the default workqueue limiter without disruption.E2E Scenario Test Plan
Add an end-to-end scenario test in
tests/e2e/testdata/scenarios/backoff_max_delay_zero/script.yamlusing a lightweight direct resource (such asLoggingLogMetric). Userecord-real-gcpskill to record real GCP logs:backoff-max-delay-in-seconds: "0":Apply a
LoggingLogMetricwith an invalid filter query andbackoff-max-delay-in-seconds: "0".Ready=False(reason: UpdateFailed).TEST: APPLY-NO-WAIT):Wait 10–15 seconds without updating the resource.
reconcile.Result{}, nilprevented workqueue re-enqueuing..spec& Immediate Recovery (TEST: WAIT-FOR-HTTP-REQUEST):Update
.specwith a valid filter expression.Ready=True.TEST: DELETE):Delete the resource and verify finalizer removal and clean deletion.
Acceptance Criteria
pkg/controller/ratelimiter/dynamicratelimiter.goimplemented withParseBackoffMaxDelayandDynamicRateLimiter.backoff-max-delay-in-seconds: "0"logs the halt, recordsReady=Falsein.status.conditions, and returnsreconcile.Result{}, nil.backoff-max-delay-in-seconds: "0"does not re-enqueue after 10–20 minutes automatically.backoff-max-delay-in-seconds: "<seconds>"(e.g.,"600") caps exponential retry delay at the specified ceiling..specimmediately interrupts any pending backoff delay and triggers reconciliation.deletionTimestamp != nil) bypass error retry delays.pkg/controller/ratelimiter/dynamicratelimiter_test.goandpkg/controller/direct/directbase/directbase_controller_test.gocovering"0", positive integers, malformed input, success resets, and deletion bypass.tests/e2e/testdata/scenarios/backoff_max_delay_zero/passing against both real GCP and MockGCP.