Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
56 changes: 53 additions & 3 deletions docs/operations/deployment.md
Original file line number Diff line number Diff line change
Expand Up @@ -79,7 +79,9 @@ Key chart values:
| `manager.image.digest` | `""` | Pin the controller image by digest. Wins over `tag`; the only form that names the exact bytes you verified |
| `manager.imagePullSecrets` | `[]` | Pull secrets for the controller image |
| `manager.resources` | `10m/500m` CPU, `1Gi/1Gi` memory | Controller resource requests/limits |
| `manager.affinity` | `{}` | Controller pod affinity |
| `manager.affinity` | `{}` | Replace the complete controller affinity; when empty, the chart prefers spreading replicas across nodes |
| `pdb.enabled` | `false` | Create a controller PodDisruptionBudget |
| `pdb.minAvailable` | `1` | Minimum ready controller pods during voluntary eviction; integer or percentage |
| `metrics.port` | `8443` | Controller metrics port |
| `metrics.serviceMonitor.enabled` | `true` | Install a `ServiceMonitor` (requires the Prometheus Operator CRDs; set to `false` on clusters without them) |

Expand Down Expand Up @@ -151,7 +153,55 @@ manager:
replicas: 2
```

Only one replica holds the leader lease at a time. Standby replicas take over automatically if the leader fails. No additional configuration is required.
Only one replica holds the leader lease at a time. Standby replicas take over automatically if the leader fails, with a temporary reconciliation gap while leadership is acquired. Existing workloads continue independently if their nodes remain healthy.

### Node maintenance and disruption protection

To retain a ready controller replica during voluntary evictions such as `kubectl drain`, enable the optional PodDisruptionBudget (PDB). `nvcrectl setup init` currently has no PDB configuration flags, so manage these values through a direct Helm install or upgrade, or through your GitOps values.

```yaml
manager:
replicas: 2
pdb:
enabled: true
minAvailable: 1
```

By default, the chart gives controller replicas a preferred pod anti-affinity rule for `kubernetes.io/hostname`. The scheduler spreads replicas across nodes when possible but may co-locate them when necessary. A non-empty `manager.affinity` replaces that complete default; it is not merged with the preferred rule.

**Upgrade note:** Earlier chart versions rendered no affinity when `manager.affinity` was empty. Upgrading from those versions with empty `manager.affinity` adds preferred hostname anti-affinity and triggers a controller Deployment rollout, even if `pdb.enabled` is false. With multiple replicas, the scheduler prefers placing them on different nodes, but still permits co-location and scheduling on a single-node cluster.

For a hard HA guarantee, include both infrastructure-node placement and required pod anti-affinity in the override. The example below assumes infrastructure nodes carry the `node-role.kubernetes.io/infra` label; replace that key with the label used by your cluster. Set `app.kubernetes.io/instance` to your Helm release name and `app.kubernetes.io/name` to the chart's rendered name label (`nvcre` by default; adjust it if you change `nameOverride`). Both should match the controller Deployment's selector:

```yaml
manager:
replicas: 2
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: node-role.kubernetes.io/infra
operator: Exists
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchLabels:
app.kubernetes.io/name: nvcre
app.kubernetes.io/instance: nvcre
control-plane: manager
topologyKey: kubernetes.io/hostname
```

The default preferred rule does not guarantee separation, so both replicas may still share a node. During a drain, one pod can be evicted, but the PDB then waits for the Deployment to schedule and ready a replacement elsewhere before allowing the other eviction. If no replacement can become Ready, the drain remains blocked. Co-location also leaves both replicas exposed to an involuntary failure of that node, which a PDB cannot prevent.

This placement requires at least two eligible infrastructure nodes. For rolling upgrades with two replicas, provide a third eligible node with capacity for a controller pod. The chart does not set a Deployment strategy, so the Kubernetes rolling-update defaults allow one surge pod and zero unavailable pods at this replica count. On exactly two nodes, the required anti-affinity also matches the old replicas, leaving the surge pod Pending and the rollout stalled. If only two nodes are available, retain preferred anti-affinity, or customize the Deployment's `spec.strategy.rollingUpdate.maxUnavailable` to `1` through your deployment tooling. The chart does not expose this strategy setting as a Helm value; allowing one unavailable replica also temporarily reduces redundancy during upgrades.

The PDB can still allow eviction of the leader; it preserves a ready replica, not uninterrupted reconciliation or process memory. It does not protect against node failure, direct pod deletion, or Deployment rolling updates.

With **one replica and `minAvailable: 1`, the PDB blocks node drains even when no Certification is running**. Before maintenance, either scale to two and wait for a ready replica on another node, or arrange a maintenance window, temporarily disable the PDB (or set `minAvailable: 0`), and restore protection after the replacement is ready. Completing a Certification does not automatically relax the budget. The PDB is disabled by default to preserve existing maintenance behavior.

More generally, voluntary eviction of a healthy controller pod is blocked while the current ready replica count is at or below the required minimum. For example, two replicas with `minAvailable: 2` or `minAvailable: "100%"` leave no room for voluntary eviction. Percentage minimums are calculated from the desired replica count and rounded up to a whole number of pods.

## RBAC requirements

Expand Down Expand Up @@ -362,7 +412,7 @@ Use this checklist before going live. Each item addresses a specific risk surfac
| **Network policy** | Required | Restrict egress to the Kubernetes API server and DNS only. No NetworkPolicy ships with the chart — add one for your environment. |
| **RBAC audit** | Required | Run `kubectl get clusterrole nvcre-manager-role -o yaml` and verify the permissions match your security requirements. |
| **TLS for metrics** | Recommended | The default ServiceMonitor uses `insecureSkipVerify: true`. Configure cert-manager to issue a serving certificate for the controller's metrics endpoint. |
| **Controller node affinity** | Recommended | Schedule the controller on infrastructure nodes, not GPU nodes, using the `manager.affinity` chart value to avoid consuming GPU resources. |
| **Controller node affinity** | Recommended | Schedule the controller on infrastructure nodes, not GPU nodes, using `manager.affinity`. This value replaces the chart's complete default affinity, so include pod anti-affinity in the override when running multiple replicas. |
| **Image provenance** | Recommended | Verify the image signature and its SLSA provenance against the exact signing identity before deploying, then pin what you verified with `--set manager.image.digest=sha256:...` rather than deploying by tag — a tag can be repointed after you check it. The provenance names the commit, ref and workflow that built it. See [Verifying release artifacts](./verifying-artifacts.md). Scan images with your vulnerability tooling before deployment. |
| **Pod Security Standards** | Verify | The controller runs as non-root with `seccompProfile: RuntimeDefault`, a read-only root filesystem, and all capabilities dropped. Verify with `kubectl get pod -n nvcre -o yaml`. |
| **CRD backup** | Recommended | Include the NVCRE CRDs in your cluster backup strategy. Certification resources contain node health state that may be needed for audit. |
Expand Down
15 changes: 13 additions & 2 deletions helm/cluster-readiness-engine/templates/deployment.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -23,9 +23,20 @@ spec:
{{- include "nvcre.labels" . | nindent 8 }}
control-plane: manager
spec:
{{- with .Values.manager.affinity }}
{{- if .Values.manager.affinity }}
affinity:
{{- toYaml . | nindent 8 }}
{{- toYaml .Values.manager.affinity | nindent 8 }}
{{- else }}
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels:
{{- include "nvcre.selectorLabels" . | nindent 20 }}
control-plane: manager
topologyKey: kubernetes.io/hostname
{{- end }}
containers:
- args:
Expand Down
18 changes: 18 additions & 0 deletions helm/cluster-readiness-engine/templates/pdb.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0

{{- if .Values.pdb.enabled }}
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: {{ include "nvcre.resourceName" (dict "suffix" "manager" "context" $) }}
namespace: {{ .Release.Namespace }}
labels:
{{- include "nvcre.labels" . | nindent 4 }}
spec:
minAvailable: {{ .Values.pdb.minAvailable | toYaml }}
selector:
matchLabels:
{{- include "nvcre.selectorLabels" . | nindent 6 }}
control-plane: manager
{{- end }}
9 changes: 9 additions & 0 deletions helm/cluster-readiness-engine/values.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -29,11 +29,20 @@ manager:
requests:
cpu: 10m
memory: 1Gi
# Replaces the complete default affinity when non-empty. Include pod anti-affinity
# explicitly when combining custom node placement with multiple replicas.
affinity: {}
tolerations:
- operator: Exists
terminationGracePeriodSeconds: 10

# Protect the controller from voluntary evictions. With one replica and
# minAvailable: 1, node drains are blocked even when no Certification is running.
pdb:
enabled: false
# Minimum ready controller pods, as an integer or percentage (for example "50%").
minAvailable: 1

metrics:
port: 8443
serviceMonitor:
Expand Down
196 changes: 196 additions & 0 deletions test/helm/pdb_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,196 @@
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
// SPDX-License-Identifier: Apache-2.0

package helm

import (
"bytes"
"io"
"reflect"
"testing"

appsv1 "k8s.io/api/apps/v1"
policyv1 "k8s.io/api/policy/v1"
"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
"k8s.io/apimachinery/pkg/runtime"
"k8s.io/apimachinery/pkg/util/intstr"
"k8s.io/apimachinery/pkg/util/yaml"
)

func TestHelmTemplateControllerPDB(t *testing.T) {
requireHelm(t)
const enablePDB = "pdb.enabled=true"
dir := chartDir(t)
requireChartInputs(t, dir)
for _, tc := range []struct {
name string
set []string
want *intstr.IntOrString
}{
{name: "disabled by default"},
{name: "explicitly disabled", set: []string{"pdb.enabled=false"}},
{
name: "singleton protection",
set: []string{enablePDB},
want: new(intstr.FromInt32(1)),
},
{
name: "two replicas",
set: []string{enablePDB, "manager.replicas=2"},
want: new(intstr.FromInt32(1)),
},
{
name: "integer override",
set: []string{enablePDB, "pdb.minAvailable=2", "nameOverride=custom"},
want: new(intstr.FromInt32(2)),
},
{
name: "zero permits maintenance",
set: []string{enablePDB, "pdb.minAvailable=0"},
want: new(intstr.FromInt32(0)),
},
{
name: "percentage",
set: []string{enablePDB, "pdb.minAvailable=50%"},
want: new(intstr.FromString("50%")),
},
} {
t.Run(tc.name, func(t *testing.T) {
rendered, err := helmTemplate(dir, tc.set)
if err != nil {
t.Fatal(err)
}
var pdbs []policyv1.PodDisruptionBudget
var deployment appsv1.Deployment
decoder := yaml.NewYAMLOrJSONDecoder(bytes.NewReader(rendered), 4096)
for {
var obj unstructured.Unstructured
if err := decoder.Decode(&obj); err != nil {
if err == io.EOF {
break
}
t.Fatal(err)
}
switch obj.GetKind() {
case "PodDisruptionBudget":
var pdb policyv1.PodDisruptionBudget
if err := runtime.DefaultUnstructuredConverter.FromUnstructured(obj.Object, &pdb); err != nil {
t.Fatal(err)
}
if obj.GetAPIVersion() != "policy/v1" {
t.Fatalf("unexpected API version: %s", obj.GetAPIVersion())
}
pdbs = append(pdbs, pdb)
case deploymentKind:
if err := runtime.DefaultUnstructuredConverter.FromUnstructured(obj.Object, &deployment); err != nil {
t.Fatal(err)
}
}
}
if tc.want == nil {
if len(pdbs) != 0 {
t.Fatal("disabled PDB was rendered")
}
return
}
if len(pdbs) != 1 {
t.Fatalf("got %d PDBs, want 1", len(pdbs))
}
pdb := pdbs[0]
if !reflect.DeepEqual(pdb.Spec.MinAvailable, tc.want) {
t.Fatalf("minAvailable = %v, want %v", pdb.Spec.MinAvailable, tc.want)
}
if pdb.Spec.MaxUnavailable != nil {
t.Fatal("maxUnavailable must be unset")
}
if deployment.Name == "" || pdb.Name != deployment.Name || pdb.Namespace != deployment.Namespace {
t.Fatal("PDB identity does not match manager Deployment")
}
if pdb.Spec.Selector == nil || len(pdb.Spec.Selector.MatchLabels) == 0 ||
!reflect.DeepEqual(pdb.Spec.Selector, deployment.Spec.Selector) {
t.Fatal("PDB selector does not match manager Deployment")
}
})
}
}

func TestHelmTemplateDefaultControllerAntiAffinity(t *testing.T) {
requireHelm(t)
dir := chartDir(t)
requireChartInputs(t, dir)

rendered, err := helmTemplate(dir, nil)
if err != nil {
t.Fatal(err)
}

deployment := decodeManagerDeployment(t, rendered)
affinity := deployment.Spec.Template.Spec.Affinity
if affinity == nil || affinity.PodAntiAffinity == nil {
t.Fatal("default pod anti-affinity was not rendered")
}
preferred := affinity.PodAntiAffinity.PreferredDuringSchedulingIgnoredDuringExecution
if len(preferred) != 1 {
t.Fatalf("got %d preferred pod anti-affinity terms, want 1", len(preferred))
}
term := preferred[0]
if term.Weight != 100 {
t.Fatalf("anti-affinity weight = %d, want 100", term.Weight)
}
if term.PodAffinityTerm.TopologyKey != "kubernetes.io/hostname" {
t.Fatalf("anti-affinity topology key = %q, want kubernetes.io/hostname",
term.PodAffinityTerm.TopologyKey)
}
if !reflect.DeepEqual(term.PodAffinityTerm.LabelSelector, deployment.Spec.Selector) {
t.Fatal("anti-affinity selector does not match manager Deployment")
}
}

func TestHelmTemplateControllerAffinityOverrideReplacesDefault(t *testing.T) {
requireHelm(t)
dir := chartDir(t)
requireChartInputs(t, dir)

const requiredNodeAffinity = "manager.affinity.nodeAffinity.requiredDuringSchedulingIgnoredDuringExecution."
rendered, err := helmTemplate(dir, []string{
requiredNodeAffinity + "nodeSelectorTerms[0].matchExpressions[0].key=node-role.kubernetes.io/control-plane",
requiredNodeAffinity + "nodeSelectorTerms[0].matchExpressions[0].operator=Exists",
})
if err != nil {
t.Fatal(err)
}

deployment := decodeManagerDeployment(t, rendered)
affinity := deployment.Spec.Template.Spec.Affinity
if affinity == nil || affinity.NodeAffinity == nil {
t.Fatal("custom node affinity was not rendered")
}
if affinity.PodAntiAffinity != nil {
t.Fatal("default pod anti-affinity was retained with a custom affinity override")
}
}

func decodeManagerDeployment(t *testing.T, rendered []byte) appsv1.Deployment {
t.Helper()
decoder := yaml.NewYAMLOrJSONDecoder(bytes.NewReader(rendered), 4096)
for {
var obj unstructured.Unstructured
if err := decoder.Decode(&obj); err != nil {
if err == io.EOF {
break
}
t.Fatal(err)
}
if obj.GetKind() != deploymentKind {
continue
}

var deployment appsv1.Deployment
if err := runtime.DefaultUnstructuredConverter.FromUnstructured(obj.Object, &deployment); err != nil {
t.Fatal(err)
}
return deployment
}
t.Fatal("rendered chart has no manager Deployment")
return appsv1.Deployment{}
}
6 changes: 4 additions & 2 deletions test/helm/render_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,8 @@ import (

const helmTemplateTimeout = 30 * time.Second

const deploymentKind = "Deployment"

func TestHelmTemplateRendersConcurrencyArgs(t *testing.T) {
requireHelm(t)
chartDir := chartDir(t)
Expand Down Expand Up @@ -133,7 +135,7 @@ func managerArgs(rendered []byte) ([]string, error) {
if err != nil {
return nil, fmt.Errorf("decode helm template output: %w", err)
}
if obj.GetKind() != "Deployment" {
if obj.GetKind() != deploymentKind {
continue
}

Expand Down Expand Up @@ -162,7 +164,7 @@ func managerImage(rendered []byte) (string, error) {
if err != nil {
return "", fmt.Errorf("decode helm template output: %w", err)
}
if obj.GetKind() != "Deployment" {
if obj.GetKind() != deploymentKind {
continue
}
var dep appsv1.Deployment
Expand Down
Loading