Skip to content

[V2] Executor livelocks on a gracefully-deleted pod until the task's retry budget is consumed #8066

Description

@Pankrat

Summary

When a task pod is deleted gracefully while its container does not exit immediately, the
TaskAction reconciler enters a livelock: it alternates between reporting
UnexpectedObjectDeletion and re-creating a pod under the same name that already exists, at
roughly 12 iterations per second, for as long as the pod object survives. MaxSystemFailures
cannot bound it, because a counter reset in the loop's other half clears it every lap.

Each iteration issues a Create that fails with AlreadyExists, and each of those failed creates
leaks one pod's worth of ResourceQuota usage (a known Kubernetes behaviour — admission charges
before the registry create and does not roll back: kubernetes/kubernetes#70563,
kubernetes/kubernetes#106842, kubernetes/kubernetes#110080). In a namespace with a pods quota,
the loop therefore exhausts the quota in seconds. The task's own next attempt is then rejected with
403 Forbidden, which the executor classifies as a user error and retries with no backoff,
burning every remaining attempt in a few hundred milliseconds without ever creating a pod.

Net effect: a routine graceful pod deletion kills the task and temporarily exhausts a quota shared
by every other workload in the namespace.

This is a regression against v1 propeller, which has the same UnexpectedObjectDeletion branch but
recovers differently (see "How v1 handled the same situation").

Versions

Observed on v2.0.35 (cr.flyte.org/flyteorg/flyte-binary-v2:v2.0.35), client flyte 2.5.16.
Every code path cited below is present in v2.0.49; line numbers are from that tag.

Reproduction

A pods ResourceQuota in the task namespace (any modest cap):

apiVersion: v1
kind: ResourceQuota
metadata:
  name: repro-pod-cap
  namespace: flyte
spec:
  hard:
    pods: "100"

A task that does not exit promptly on SIGTERM, with retries=5 so the budget burn is visible:

import asyncio
import signal

import flyte

env = flyte.TaskEnvironment(name="preemption-repro")


@env.task(retries=5)
async def sleeps_through_sigterm() -> str:
    signal.signal(signal.SIGTERM, signal.SIG_IGN)
    print("running", flush=True)
    await asyncio.sleep(3600)
    return "done"


if __name__ == "__main__":
    flyte.init_from_config()
    print(flyte.run(sleeps_through_sigterm).url)

Then, once the pod is Running:

kubectl -n flyte delete pod <run>-a0-0

kubectl -n flyte logs -f deploy/<flyte-release> -c flyte
kubectl -n flyte get resourcequota repro-pod-cap -w -o custom-columns='PODS:.status.used.pods'

A plain kubectl delete pod is enough - it is an ordinary graceful delete, the same call a
scheduler makes when it preempts. We hit this with NVIDIA KAI; kube-scheduler preemption,
cluster-autoscaler scale-down and node drain all take the same path.

The explicit SIG_IGN is for determinism, not necessity: in our runs a task with no signal
handling at all also held its pod Running with a deletionTimestamp for the full 30 s grace
period, so the default case is already affected. Ignoring the signal just removes any dependency on
how the runtime happens to handle it.

Observed: ~90 UnexpectedObjectDeletion errors in 8s, used.pods climbing to 100 in lockstep
(one increment per ~80 ms), then attempts 2–6 in 300 ms, then the action FAILED. Raising
terminationGracePeriodSeconds scales the leak linearly.

Mechanism

All line numbers are v2.0.49.

1. A graceful delete is read as an unexpected one. executor/pkg/plugin/k8s/plugin_manager.go:219

if !p.Phase().IsTerminal() && o.GetDeletionTimestamp() != nil {
    // -> PhaseInfoSystemRetryableFailure("UnexpectedObjectDeletion", ...)

A pod under a graceful delete has a deletionTimestamp and remains non-terminal for its whole
grace period, so this branch is true continuously for up to
terminationGracePeriodSeconds — it is not a one-shot condition.

2. The recovery re-creates a pod under the same name, and misreports success.
executor/pkg/controller/taskaction_controller.go:812-814 routes the system-retryable failure to
resetPluginResource (:501), which aborts and clears Status.PluginState but keeps
Status.Attempts
. The next reconcile therefore takes the PluginPhaseNotStarted path into
launchResource with the same generated name. There, plugin_manager.go:117-143:

err = pm.kubeClient.GetClient().Create(ctx, o)
if err != nil && !k8serrors.IsAlreadyExists(err) {   // :118  AlreadyExists is swallowed
    ...
}
return pluginsCore.DoTransition(pluginsCore.PhaseInfoQueued(..., "task submitted to K8s")), nil  // :143

The still-terminating pod makes Create return AlreadyExists, which is swallowed, and the
transition reports Queued as though a fresh pod had been submitted.

The Abort inside resetPluginResource does not help: it issues a plain Delete with no grace
override (plugin_manager.go:465), and a second graceful delete of an already-terminating pod is a
no-op.

3. MaxSystemFailures can never bound the loop. taskaction_controller.go:827

taskAction.Status.SystemFailures = 0   // reset on ANY non-system transition

This sits after the early return at :812, so the loop alternates
SystemFailures++ (:565) → Queued → reset → SystemFailures++ → … and never reaches
DefaultMaxSystemFailures = 3 (:69).

The loop is self-triggering: the controller is registered with For(&TaskAction{}) and no
predicate (:1420-1423), and both halves of the loop write TaskAction status
(recordSystemError persists the counter, the Queued half persists PluginState). Each write is
a watch event that re-enqueues the same object, so the cadence is API latency, not the requeue
duration (which is why the observed ~12 Hz is far above the 10 s default requeue, and why making
the requeue configurable in #7770 does not affect this).

4. The resulting quota rejection is charged to the user, with no backoff.
plugin_manager.go:119-121

if k8serrors.IsForbidden(err) {
    return pluginsCore.DoTransition(pluginsCore.PhaseInfoRetryableFailure("RuntimeFailure", err.Error(), nil)), nil
}

A ResourceQuota rejection is 403 Forbidden, so it becomes a USER-kind retryable failure. The
in-place restart path (taskaction_controller.go:832-847) increments Status.Attempts and
re-enters immediately — there is no backoff on retryable failures anywhere in this path — so the
whole budget is spent at API speed. The action's final error_info is:

pods "<run>-a0-5" is forbidden: exceeded quota: namespace-defaults,
requested: pods=1, used: pods=100, limited: pods=100
kind: KIND_USER

No pod was created for attempts 2–6; there is no pod object or event for any of them.

How v1 handled the same situation

v1 propeller has the identical deletion branch
(flytepropeller/pkg/controller/nodes/task/k8s/plugin_manager.go:346-352 on master), with a
comment naming the intended cases: a vanished kubelet leaving pods stuck behind the flyte
finalizer, and "when a user deletes a Pod directly". Two things made it harmless there:

  • A SYSTEM retryable failure started a new attempt (nodes/executor.go:830,
    isEligibleForRetry), so the next pod had a new name and AlreadyExists could never collide
    with the dying one. v2's resetPluginResource reuses the attempt number, which is the root of
    the livelock.
  • exceeded quota on Create was recognised explicitly and mapped to WaitingForResources
    behind a backoff controller (plugin_manager.go:247-258 on master,
    backoff.IsResourceQuotaExceeded). v2 dropped that and lets the generic IsForbidden branch
    turn it into a user error.

Why this matters beyond one cluster

  • Any graceful pod deletion triggers it: scheduler preemption (kube-scheduler, Volcano, KAI), node
    drain, cluster-autoscaler scale-down, a manual kubectl delete pod.
  • The blast radius is the whole namespace, not the task: while the quota is exhausted, every
    workload's pod creation in that namespace is rejected.
  • It penalises correct behaviour. A task that handles SIGTERM to checkpoint before exiting — the
    thing users are told to do on preemptible capacity — holds its pod object longer and therefore
    leaks more quota. A task that ignores shutdown entirely and gets SIGKILLed is affected least.
    A trainer that checkpoints for 60 s leaks ~700 charges.

Proposed fixes, in order of what they buy

  1. Treat a non-terminal pod with a deletionTimestamp as "still shutting down", not as a
    failure
    (plugin_manager.go:219). A graceful delete is a normal, time-bounded state. Keep
    observing until the object is gone — the NotFound path already produces
    ResourceDeletedExternally (:164) — and only then decide what to do. Alternatively, restore
    the v1 semantics and start a new attempt (new name) on a SYSTEM failure instead of
    re-launching under the same name. Either one removes the cause: no Create is ever issued
    against a name that is known to be occupied, so nothing leaks. This is the fix that matters.
  2. Do not classify an admission/quota 403 as a user error, and add backoff to retryable
    failures
    (plugin_manager.go:119, taskaction_controller.go:832-847). A ResourceQuota
    rejection is a transient platform condition; v1's WaitingForResources + backoff is the prior
    art. Even with fix 1, an unrelated quota exhaustion still burns a task's whole budget in
    milliseconds today.
  3. launchResource should not swallow AlreadyExists into Queued (:118). It reports
    "submitted" for a pod it did not submit. Note this is hardening, not a fix on its own: the leak
    is per failed Create, and the recovery would still issue one per lap.
  4. Do not reset Status.SystemFailures on a transition produced by the same recovery cycle
    (taskaction_controller.go:827). Also hardening only — and not safe alone: bounding the
    loop at 3 without fix 1 turns every graceful deletion into a permanent
    MaxSystemFailuresExceeded failure within a second.

Happy to open a PR for (1) if the maintainers agree on the direction (observe-until-gone vs.
new-attempt).

Related: #8067 - with the DisruptionTarget condition present (set by KAI before the
delete, so it is on the pod for the whole grace window), the executor can also tell "preempted,
shutting down" from "deleted behind my back" in this very branch.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions