Summary
The gcm Helm chart's health-checks DaemonSet assumes the NVIDIA driver toolkit lives at the conventional /usr/local/nvidia path and that a DCGM host engine is already running on the node. Neither holds on Google Kubernetes Engine (GKE) nodes running Container-Optimized OS (COS). As shipped, the health-checks pods come up but check-nvidia-smi and check-dcgmi fail on every GPU node.
We currently make GCM work on GKE with a post-install kubectl patch against the rendered DaemonSet (full script at the bottom). That works but is fragile: it re-mutates a Helm-managed object by container index, so every helm upgrade reverts it and any template change upstream can silently break it.
This issue proposes adding a small set of generic, opt-in chart knobs so GKE (and any cluster with a non-standard driver layout) can be supported purely through a values file, with no GKE-specific branching inside the templates.
Verification status: Items 1 and 2 below are reproduced and well understood. Items 3 and 4 are based on what our working patch does and our reading of GKE behavior, but we have not independently confirmed the exact root cause yet (see the per-item notes). Please don't treat them as hard blockers, they may turn out to be conventions rather than strict requirements. Happy to dig in further if a maintainer wants confirmation before any of this lands.
Environment
- GKE node pool on Container-Optimized OS (COS) with the GKE-managed NVIDIA driver installer.
gcm chart installed via oci://ghcr.io/facebookresearch/charts/gcm.
healthChecks.enabled: true.
What breaks on GKE COS, and why
-
Driver path differs. GKE installs the driver at /home/kubernetes/bin/nvidia/ (libs under lib64/, binaries under bin/), not /usr/local/nvidia/. With chart defaults, libnvidia-ml.so and nvidia-smi/dcgmi are not on LD_LIBRARY_PATH/PATH, so check-nvidia-smi and check-dcgmi cannot find them. The DaemonSet already mounts the host root at /host, so the files are reachable, just not on the search paths. (Confirmed.)
-
No DCGM host engine running. GKE does not pre-start nv-hostengine, so check-dcgmi has nothing to talk to. It needs to be started inside the pod before NPD runs. (Confirmed.)
-
NPD Prometheus port may collide. GKE runs its own node-problem-detector. Our patch moves GCM's NPD Prometheus port to 20357 to avoid the default. Not independently confirmed: this is what our working config does, but we haven't verified the default port is actually contended on GKE vs. simply a value we chose to be safe. If it's not a real collision, this just becomes an optional override.
-
DCGM diagnostic plugin / persistence-mode setup. Our patch symlinks the cudaless DCGM plugin (so diagnostics run without a full CUDA install) and enables persistence mode (nvidia-smi -pm 1, to reduce per-check latency). Not strictly GKE-specific: these are generally useful and likely belong behind an optional knob rather than being GKE-gated.
Why this can't be expressed with the chart today
healthChecks.extraEnv already exists, so item 1 (the LD_LIBRARY_PATH / PATH overrides) is solvable with values today. But:
- The container
command/args are hardcoded in templates/health-checks-daemonset.yaml, so there is no way to run init logic (start nv-hostengine, symlink the plugin, enable persistence mode) before NPD execs.
- The Prometheus port is baked into the args rather than read from a value, so item 3 can't be set without rewriting the whole
args.
- There are no
extraVolumes / extraVolumeMounts passthroughs, so the host driver directory can't be surfaced at the path standard tooling expects.
Proposed chart changes (all backward-compatible, all opt-in)
I'd like maintainer guidance on shape before opening a PR, but the proposal is:
-
healthChecks.initScript (string, default ""). When set, the template renders the container as command: ["/bin/sh","-c"] and args: ["<initScript>\nexec /usr/local/bin/node-problem-detector ..."]. When empty, behaviour is exactly as today. This carries the nv-hostengine start, the plugin symlink, and persistence mode without any if eq .Values.platform "gke" in the chart.
-
healthChecks.prometheus.port — already present; ensure the rendered --prometheus-port reads from it (so a port override is possible) instead of being a constant.
-
healthChecks.extraVolumes / healthChecks.extraVolumeMounts (lists, default []) — standard passthroughs so the host driver lib dir can be mounted at /usr/local/nvidia/lib64. Common pattern in other charts.
-
healthChecks.extraEnv — already exists; no change. Used for the LD_LIBRARY_PATH / PATH overrides.
-
Docs + example values file (examples/values-gke.yaml) wiring the above together as the first worked platform example.
Why knobs, not a platform: gke switch
Keeping the chart platform-agnostic and expressing GKE entirely as a values file avoids a combinatorial gke/eks/aks/on-prem branch matrix inside the templates. The only genuinely new mechanisms requested are initScript and the extraVolumes/extraVolumeMounts passthroughs, both reusable by any cluster with a non-standard driver layout, not just GKE.
Example: proposed GKE values file
# examples/values-gke.yaml
healthChecks:
enabled: true
prometheus:
port: 20357 # avoid GKE's own node-problem-detector (see note in issue)
extraEnv:
- name: LD_LIBRARY_PATH
value: /usr/local/nvidia/lib64:/host/home/kubernetes/bin/nvidia/lib64
- name: PATH
value: /usr/local/nvidia/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/host/home/kubernetes/bin/nvidia/bin
extraVolumes:
- name: nvidia-libs
hostPath:
path: /home/kubernetes/bin/nvidia/lib64
type: DirectoryOrCreate
extraVolumeMounts:
- name: nvidia-libs
mountPath: /usr/local/nvidia/lib64
readOnly: true
initScript: |
NVIDIA_LIB=/host/home/kubernetes/bin/nvidia/lib64
if [ -f "${NVIDIA_LIB}/libnvidia-ml.so" ]; then
nv-hostengine -n -b ALL >/dev/null 2>&1 &
ln -sf /usr/libexec/datacenter-gpu-manager-4/plugins/cudaless \
/usr/libexec/datacenter-gpu-manager-4/plugins/cuda13 2>/dev/null
/host/home/kubernetes/bin/nvidia/bin/nvidia-smi -pm 1 >/dev/null 2>&1
sleep 2
echo "GPU node: nv-hostengine started, persistence mode enabled"
fi
With these knobs, deployment becomes:
helm install gcm oci://ghcr.io/facebookresearch/charts/gcm -f examples/values-gke.yaml
and survives helm upgrade with no post-install patching.
Current workaround (post-install patch we run today)
For reference, this is the kubectl patch we apply after helm install. The goal of this issue is to make this unnecessary.
kubectl patch ds gcm-health-checks -n kube-system --type=json -p='[
{"op":"add","path":"/spec/template/spec/containers/0/env/-",
"value":{"name":"LD_LIBRARY_PATH",
"value":"/usr/local/nvidia/lib64:/host/home/kubernetes/bin/nvidia/lib64"}},
{"op":"add","path":"/spec/template/spec/containers/0/env/-",
"value":{"name":"PATH",
"value":"/usr/local/nvidia/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/host/home/kubernetes/bin/nvidia/bin"}},
{"op":"replace","path":"/spec/template/spec/containers/0/command",
"value":["/bin/sh","-c"]},
{"op":"replace","path":"/spec/template/spec/containers/0/args",
"value":["NVIDIA_LIB=/host/home/kubernetes/bin/nvidia/lib64; if [ -f ${NVIDIA_LIB}/libnvidia-ml.so ]; then nv-hostengine -n -b ALL >/dev/null 2>&1 & ln -sf /usr/libexec/datacenter-gpu-manager-4/plugins/cudaless /usr/libexec/datacenter-gpu-manager-4/plugins/cuda13 2>/dev/null; /host/home/kubernetes/bin/nvidia/bin/nvidia-smi -pm 1 >/dev/null 2>&1; sleep 2; echo \"GPU node: nv-hostengine started, persistence mode enabled\"; fi; exec /usr/local/bin/node-problem-detector --logtostderr --config.custom-plugin-monitor=/config/gcm_monitor.json --prometheus-address=0.0.0.0 --prometheus-port=20357 --port=0"]},
{"op":"add","path":"/spec/template/spec/volumes/-",
"value":{"name":"nvidia-libs","hostPath":{"path":"/home/kubernetes/bin/nvidia/lib64","type":"DirectoryOrCreate"}}},
{"op":"add","path":"/spec/template/spec/containers/0/volumeMounts/-",
"value":{"name":"nvidia-libs","mountPath":"/usr/local/nvidia/lib64","readOnly":true}}
]'
Questions for maintainers
- Are you open to platform-portability knobs of this shape, or would you prefer a different mechanism (e.g. an init container instead of an inline
initScript)?
- Is
examples/values-gke.yaml + a docs section the right home for the GKE recipe, or do you want platform recipes elsewhere?
- Does a change of this size need a formal RFC, or is a PR with the above sufficient?
Happy to open the PR (CLA signed) once the approach is blessed.
Summary
The
gcmHelm chart'shealth-checksDaemonSet assumes the NVIDIA driver toolkit lives at the conventional/usr/local/nvidiapath and that a DCGM host engine is already running on the node. Neither holds on Google Kubernetes Engine (GKE) nodes running Container-Optimized OS (COS). As shipped, the health-checks pods come up butcheck-nvidia-smiandcheck-dcgmifail on every GPU node.We currently make GCM work on GKE with a post-install
kubectl patchagainst the rendered DaemonSet (full script at the bottom). That works but is fragile: it re-mutates a Helm-managed object by container index, so everyhelm upgradereverts it and any template change upstream can silently break it.This issue proposes adding a small set of generic, opt-in chart knobs so GKE (and any cluster with a non-standard driver layout) can be supported purely through a values file, with no GKE-specific branching inside the templates.
Environment
gcmchart installed viaoci://ghcr.io/facebookresearch/charts/gcm.healthChecks.enabled: true.What breaks on GKE COS, and why
Driver path differs. GKE installs the driver at
/home/kubernetes/bin/nvidia/(libs underlib64/, binaries underbin/), not/usr/local/nvidia/. With chart defaults,libnvidia-ml.soandnvidia-smi/dcgmiare not onLD_LIBRARY_PATH/PATH, socheck-nvidia-smiandcheck-dcgmicannot find them. The DaemonSet already mounts the host root at/host, so the files are reachable, just not on the search paths. (Confirmed.)No DCGM host engine running. GKE does not pre-start
nv-hostengine, socheck-dcgmihas nothing to talk to. It needs to be started inside the pod before NPD runs. (Confirmed.)NPD Prometheus port may collide. GKE runs its own node-problem-detector. Our patch moves GCM's NPD Prometheus port to
20357to avoid the default. Not independently confirmed: this is what our working config does, but we haven't verified the default port is actually contended on GKE vs. simply a value we chose to be safe. If it's not a real collision, this just becomes an optional override.DCGM diagnostic plugin / persistence-mode setup. Our patch symlinks the
cudalessDCGM plugin (so diagnostics run without a full CUDA install) and enables persistence mode (nvidia-smi -pm 1, to reduce per-check latency). Not strictly GKE-specific: these are generally useful and likely belong behind an optional knob rather than being GKE-gated.Why this can't be expressed with the chart today
healthChecks.extraEnvalready exists, so item 1 (theLD_LIBRARY_PATH/PATHoverrides) is solvable with values today. But:command/argsare hardcoded intemplates/health-checks-daemonset.yaml, so there is no way to run init logic (startnv-hostengine, symlink the plugin, enable persistence mode) before NPD execs.args.extraVolumes/extraVolumeMountspassthroughs, so the host driver directory can't be surfaced at the path standard tooling expects.Proposed chart changes (all backward-compatible, all opt-in)
I'd like maintainer guidance on shape before opening a PR, but the proposal is:
healthChecks.initScript(string, default""). When set, the template renders the container ascommand: ["/bin/sh","-c"]andargs: ["<initScript>\nexec /usr/local/bin/node-problem-detector ..."]. When empty, behaviour is exactly as today. This carries thenv-hostenginestart, the plugin symlink, and persistence mode without anyif eq .Values.platform "gke"in the chart.healthChecks.prometheus.port— already present; ensure the rendered--prometheus-portreads from it (so a port override is possible) instead of being a constant.healthChecks.extraVolumes/healthChecks.extraVolumeMounts(lists, default[]) — standard passthroughs so the host driver lib dir can be mounted at/usr/local/nvidia/lib64. Common pattern in other charts.healthChecks.extraEnv— already exists; no change. Used for theLD_LIBRARY_PATH/PATHoverrides.Docs + example values file (
examples/values-gke.yaml) wiring the above together as the first worked platform example.Why knobs, not a
platform: gkeswitchKeeping the chart platform-agnostic and expressing GKE entirely as a values file avoids a combinatorial
gke/eks/aks/on-prem branch matrix inside the templates. The only genuinely new mechanisms requested areinitScriptand theextraVolumes/extraVolumeMountspassthroughs, both reusable by any cluster with a non-standard driver layout, not just GKE.Example: proposed GKE values file
With these knobs, deployment becomes:
helm install gcm oci://ghcr.io/facebookresearch/charts/gcm -f examples/values-gke.yamland survives
helm upgradewith no post-install patching.Current workaround (post-install patch we run today)
For reference, this is the
kubectl patchwe apply afterhelm install. The goal of this issue is to make this unnecessary.Questions for maintainers
initScript)?examples/values-gke.yaml+ a docs section the right home for the GKE recipe, or do you want platform recipes elsewhere?Happy to open the PR (CLA signed) once the approach is blessed.