Skip to content

Support GKE / Container-Optimized OS deployments of the health-checks DaemonSet #163

Description

@deokcf

Summary

The gcm Helm chart's health-checks DaemonSet assumes the NVIDIA driver toolkit lives at the conventional /usr/local/nvidia path and that a DCGM host engine is already running on the node. Neither holds on Google Kubernetes Engine (GKE) nodes running Container-Optimized OS (COS). As shipped, the health-checks pods come up but check-nvidia-smi and check-dcgmi fail on every GPU node.

We currently make GCM work on GKE with a post-install kubectl patch against the rendered DaemonSet (full script at the bottom). That works but is fragile: it re-mutates a Helm-managed object by container index, so every helm upgrade reverts it and any template change upstream can silently break it.

This issue proposes adding a small set of generic, opt-in chart knobs so GKE (and any cluster with a non-standard driver layout) can be supported purely through a values file, with no GKE-specific branching inside the templates.

Verification status: Items 1 and 2 below are reproduced and well understood. Items 3 and 4 are based on what our working patch does and our reading of GKE behavior, but we have not independently confirmed the exact root cause yet (see the per-item notes). Please don't treat them as hard blockers, they may turn out to be conventions rather than strict requirements. Happy to dig in further if a maintainer wants confirmation before any of this lands.

Environment

  • GKE node pool on Container-Optimized OS (COS) with the GKE-managed NVIDIA driver installer.
  • gcm chart installed via oci://ghcr.io/facebookresearch/charts/gcm.
  • healthChecks.enabled: true.

What breaks on GKE COS, and why

  1. Driver path differs. GKE installs the driver at /home/kubernetes/bin/nvidia/ (libs under lib64/, binaries under bin/), not /usr/local/nvidia/. With chart defaults, libnvidia-ml.so and nvidia-smi/dcgmi are not on LD_LIBRARY_PATH/PATH, so check-nvidia-smi and check-dcgmi cannot find them. The DaemonSet already mounts the host root at /host, so the files are reachable, just not on the search paths. (Confirmed.)

  2. No DCGM host engine running. GKE does not pre-start nv-hostengine, so check-dcgmi has nothing to talk to. It needs to be started inside the pod before NPD runs. (Confirmed.)

  3. NPD Prometheus port may collide. GKE runs its own node-problem-detector. Our patch moves GCM's NPD Prometheus port to 20357 to avoid the default. Not independently confirmed: this is what our working config does, but we haven't verified the default port is actually contended on GKE vs. simply a value we chose to be safe. If it's not a real collision, this just becomes an optional override.

  4. DCGM diagnostic plugin / persistence-mode setup. Our patch symlinks the cudaless DCGM plugin (so diagnostics run without a full CUDA install) and enables persistence mode (nvidia-smi -pm 1, to reduce per-check latency). Not strictly GKE-specific: these are generally useful and likely belong behind an optional knob rather than being GKE-gated.

Why this can't be expressed with the chart today

healthChecks.extraEnv already exists, so item 1 (the LD_LIBRARY_PATH / PATH overrides) is solvable with values today. But:

  • The container command/args are hardcoded in templates/health-checks-daemonset.yaml, so there is no way to run init logic (start nv-hostengine, symlink the plugin, enable persistence mode) before NPD execs.
  • The Prometheus port is baked into the args rather than read from a value, so item 3 can't be set without rewriting the whole args.
  • There are no extraVolumes / extraVolumeMounts passthroughs, so the host driver directory can't be surfaced at the path standard tooling expects.

Proposed chart changes (all backward-compatible, all opt-in)

I'd like maintainer guidance on shape before opening a PR, but the proposal is:

  1. healthChecks.initScript (string, default ""). When set, the template renders the container as command: ["/bin/sh","-c"] and args: ["<initScript>\nexec /usr/local/bin/node-problem-detector ..."]. When empty, behaviour is exactly as today. This carries the nv-hostengine start, the plugin symlink, and persistence mode without any if eq .Values.platform "gke" in the chart.

  2. healthChecks.prometheus.port — already present; ensure the rendered --prometheus-port reads from it (so a port override is possible) instead of being a constant.

  3. healthChecks.extraVolumes / healthChecks.extraVolumeMounts (lists, default []) — standard passthroughs so the host driver lib dir can be mounted at /usr/local/nvidia/lib64. Common pattern in other charts.

  4. healthChecks.extraEnv — already exists; no change. Used for the LD_LIBRARY_PATH / PATH overrides.

  5. Docs + example values file (examples/values-gke.yaml) wiring the above together as the first worked platform example.

Why knobs, not a platform: gke switch

Keeping the chart platform-agnostic and expressing GKE entirely as a values file avoids a combinatorial gke/eks/aks/on-prem branch matrix inside the templates. The only genuinely new mechanisms requested are initScript and the extraVolumes/extraVolumeMounts passthroughs, both reusable by any cluster with a non-standard driver layout, not just GKE.

Example: proposed GKE values file

# examples/values-gke.yaml
healthChecks:
  enabled: true
  prometheus:
    port: 20357   # avoid GKE's own node-problem-detector (see note in issue)
  extraEnv:
    - name: LD_LIBRARY_PATH
      value: /usr/local/nvidia/lib64:/host/home/kubernetes/bin/nvidia/lib64
    - name: PATH
      value: /usr/local/nvidia/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/host/home/kubernetes/bin/nvidia/bin
  extraVolumes:
    - name: nvidia-libs
      hostPath:
        path: /home/kubernetes/bin/nvidia/lib64
        type: DirectoryOrCreate
  extraVolumeMounts:
    - name: nvidia-libs
      mountPath: /usr/local/nvidia/lib64
      readOnly: true
  initScript: |
    NVIDIA_LIB=/host/home/kubernetes/bin/nvidia/lib64
    if [ -f "${NVIDIA_LIB}/libnvidia-ml.so" ]; then
      nv-hostengine -n -b ALL >/dev/null 2>&1 &
      ln -sf /usr/libexec/datacenter-gpu-manager-4/plugins/cudaless \
             /usr/libexec/datacenter-gpu-manager-4/plugins/cuda13 2>/dev/null
      /host/home/kubernetes/bin/nvidia/bin/nvidia-smi -pm 1 >/dev/null 2>&1
      sleep 2
      echo "GPU node: nv-hostengine started, persistence mode enabled"
    fi

With these knobs, deployment becomes:
helm install gcm oci://ghcr.io/facebookresearch/charts/gcm -f examples/values-gke.yaml
and survives helm upgrade with no post-install patching.

Current workaround (post-install patch we run today)

For reference, this is the kubectl patch we apply after helm install. The goal of this issue is to make this unnecessary.

kubectl patch ds gcm-health-checks -n kube-system --type=json -p='[
  {"op":"add","path":"/spec/template/spec/containers/0/env/-",
   "value":{"name":"LD_LIBRARY_PATH",
            "value":"/usr/local/nvidia/lib64:/host/home/kubernetes/bin/nvidia/lib64"}},
  {"op":"add","path":"/spec/template/spec/containers/0/env/-",
   "value":{"name":"PATH",
            "value":"/usr/local/nvidia/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/host/home/kubernetes/bin/nvidia/bin"}},
  {"op":"replace","path":"/spec/template/spec/containers/0/command",
   "value":["/bin/sh","-c"]},
  {"op":"replace","path":"/spec/template/spec/containers/0/args",
   "value":["NVIDIA_LIB=/host/home/kubernetes/bin/nvidia/lib64; if [ -f ${NVIDIA_LIB}/libnvidia-ml.so ]; then nv-hostengine -n -b ALL >/dev/null 2>&1 & ln -sf /usr/libexec/datacenter-gpu-manager-4/plugins/cudaless /usr/libexec/datacenter-gpu-manager-4/plugins/cuda13 2>/dev/null; /host/home/kubernetes/bin/nvidia/bin/nvidia-smi -pm 1 >/dev/null 2>&1; sleep 2; echo \"GPU node: nv-hostengine started, persistence mode enabled\"; fi; exec /usr/local/bin/node-problem-detector --logtostderr --config.custom-plugin-monitor=/config/gcm_monitor.json --prometheus-address=0.0.0.0 --prometheus-port=20357 --port=0"]},
  {"op":"add","path":"/spec/template/spec/volumes/-",
   "value":{"name":"nvidia-libs","hostPath":{"path":"/home/kubernetes/bin/nvidia/lib64","type":"DirectoryOrCreate"}}},
  {"op":"add","path":"/spec/template/spec/containers/0/volumeMounts/-",
   "value":{"name":"nvidia-libs","mountPath":"/usr/local/nvidia/lib64","readOnly":true}}
]'

Questions for maintainers

  1. Are you open to platform-portability knobs of this shape, or would you prefer a different mechanism (e.g. an init container instead of an inline initScript)?
  2. Is examples/values-gke.yaml + a docs section the right home for the GKE recipe, or do you want platform recipes elsewhere?
  3. Does a change of this size need a formal RFC, or is a PR with the above sufficient?

Happy to open the PR (CLA signed) once the approach is blessed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions