Summary
When a task pod is deleted gracefully while its container does not exit immediately, the
TaskAction reconciler enters a livelock: it alternates between reporting
UnexpectedObjectDeletion and re-creating a pod under the same name that already exists, at
roughly 12 iterations per second, for as long as the pod object survives. MaxSystemFailures
cannot bound it, because a counter reset in the loop's other half clears it every lap.
Each iteration issues a Create that fails with AlreadyExists, and each of those failed creates
leaks one pod's worth of ResourceQuota usage (a known Kubernetes behaviour — admission charges
before the registry create and does not roll back: kubernetes/kubernetes#70563,
kubernetes/kubernetes#106842, kubernetes/kubernetes#110080). In a namespace with a pods quota,
the loop therefore exhausts the quota in seconds. The task's own next attempt is then rejected with
403 Forbidden, which the executor classifies as a user error and retries with no backoff,
burning every remaining attempt in a few hundred milliseconds without ever creating a pod.
Net effect: a routine graceful pod deletion kills the task and temporarily exhausts a quota shared
by every other workload in the namespace.
This is a regression against v1 propeller, which has the same UnexpectedObjectDeletion branch but
recovers differently (see "How v1 handled the same situation").
Versions
Observed on v2.0.35 (cr.flyte.org/flyteorg/flyte-binary-v2:v2.0.35), client flyte 2.5.16.
Every code path cited below is present in v2.0.49; line numbers are from that tag.
Reproduction
A pods ResourceQuota in the task namespace (any modest cap):
apiVersion: v1
kind: ResourceQuota
metadata:
name: repro-pod-cap
namespace: flyte
spec:
hard:
pods: "100"
A task that does not exit promptly on SIGTERM, with retries=5 so the budget burn is visible:
import asyncio
import signal
import flyte
env = flyte.TaskEnvironment(name="preemption-repro")
@env.task(retries=5)
async def sleeps_through_sigterm() -> str:
signal.signal(signal.SIGTERM, signal.SIG_IGN)
print("running", flush=True)
await asyncio.sleep(3600)
return "done"
if __name__ == "__main__":
flyte.init_from_config()
print(flyte.run(sleeps_through_sigterm).url)
Then, once the pod is Running:
kubectl -n flyte delete pod <run>-a0-0
kubectl -n flyte logs -f deploy/<flyte-release> -c flyte
kubectl -n flyte get resourcequota repro-pod-cap -w -o custom-columns='PODS:.status.used.pods'
A plain kubectl delete pod is enough - it is an ordinary graceful delete, the same call a
scheduler makes when it preempts. We hit this with NVIDIA KAI; kube-scheduler preemption,
cluster-autoscaler scale-down and node drain all take the same path.
The explicit SIG_IGN is for determinism, not necessity: in our runs a task with no signal
handling at all also held its pod Running with a deletionTimestamp for the full 30 s grace
period, so the default case is already affected. Ignoring the signal just removes any dependency on
how the runtime happens to handle it.
Observed: ~90 UnexpectedObjectDeletion errors in 8s, used.pods climbing to 100 in lockstep
(one increment per ~80 ms), then attempts 2–6 in 300 ms, then the action FAILED. Raising
terminationGracePeriodSeconds scales the leak linearly.
Mechanism
All line numbers are v2.0.49.
1. A graceful delete is read as an unexpected one. executor/pkg/plugin/k8s/plugin_manager.go:219
if !p.Phase().IsTerminal() && o.GetDeletionTimestamp() != nil {
// -> PhaseInfoSystemRetryableFailure("UnexpectedObjectDeletion", ...)
A pod under a graceful delete has a deletionTimestamp and remains non-terminal for its whole
grace period, so this branch is true continuously for up to
terminationGracePeriodSeconds — it is not a one-shot condition.
2. The recovery re-creates a pod under the same name, and misreports success.
executor/pkg/controller/taskaction_controller.go:812-814 routes the system-retryable failure to
resetPluginResource (:501), which aborts and clears Status.PluginState but keeps
Status.Attempts. The next reconcile therefore takes the PluginPhaseNotStarted path into
launchResource with the same generated name. There, plugin_manager.go:117-143:
err = pm.kubeClient.GetClient().Create(ctx, o)
if err != nil && !k8serrors.IsAlreadyExists(err) { // :118 AlreadyExists is swallowed
...
}
return pluginsCore.DoTransition(pluginsCore.PhaseInfoQueued(..., "task submitted to K8s")), nil // :143
The still-terminating pod makes Create return AlreadyExists, which is swallowed, and the
transition reports Queued as though a fresh pod had been submitted.
The Abort inside resetPluginResource does not help: it issues a plain Delete with no grace
override (plugin_manager.go:465), and a second graceful delete of an already-terminating pod is a
no-op.
3. MaxSystemFailures can never bound the loop. taskaction_controller.go:827
taskAction.Status.SystemFailures = 0 // reset on ANY non-system transition
This sits after the early return at :812, so the loop alternates
SystemFailures++ (:565) → Queued → reset → SystemFailures++ → … and never reaches
DefaultMaxSystemFailures = 3 (:69).
The loop is self-triggering: the controller is registered with For(&TaskAction{}) and no
predicate (:1420-1423), and both halves of the loop write TaskAction status
(recordSystemError persists the counter, the Queued half persists PluginState). Each write is
a watch event that re-enqueues the same object, so the cadence is API latency, not the requeue
duration (which is why the observed ~12 Hz is far above the 10 s default requeue, and why making
the requeue configurable in #7770 does not affect this).
4. The resulting quota rejection is charged to the user, with no backoff.
plugin_manager.go:119-121
if k8serrors.IsForbidden(err) {
return pluginsCore.DoTransition(pluginsCore.PhaseInfoRetryableFailure("RuntimeFailure", err.Error(), nil)), nil
}
A ResourceQuota rejection is 403 Forbidden, so it becomes a USER-kind retryable failure. The
in-place restart path (taskaction_controller.go:832-847) increments Status.Attempts and
re-enters immediately — there is no backoff on retryable failures anywhere in this path — so the
whole budget is spent at API speed. The action's final error_info is:
pods "<run>-a0-5" is forbidden: exceeded quota: namespace-defaults,
requested: pods=1, used: pods=100, limited: pods=100
kind: KIND_USER
No pod was created for attempts 2–6; there is no pod object or event for any of them.
How v1 handled the same situation
v1 propeller has the identical deletion branch
(flytepropeller/pkg/controller/nodes/task/k8s/plugin_manager.go:346-352 on master), with a
comment naming the intended cases: a vanished kubelet leaving pods stuck behind the flyte
finalizer, and "when a user deletes a Pod directly". Two things made it harmless there:
- A SYSTEM retryable failure started a new attempt (
nodes/executor.go:830,
isEligibleForRetry), so the next pod had a new name and AlreadyExists could never collide
with the dying one. v2's resetPluginResource reuses the attempt number, which is the root of
the livelock.
exceeded quota on Create was recognised explicitly and mapped to WaitingForResources
behind a backoff controller (plugin_manager.go:247-258 on master,
backoff.IsResourceQuotaExceeded). v2 dropped that and lets the generic IsForbidden branch
turn it into a user error.
Why this matters beyond one cluster
- Any graceful pod deletion triggers it: scheduler preemption (kube-scheduler, Volcano, KAI), node
drain, cluster-autoscaler scale-down, a manual kubectl delete pod.
- The blast radius is the whole namespace, not the task: while the quota is exhausted, every
workload's pod creation in that namespace is rejected.
- It penalises correct behaviour. A task that handles
SIGTERM to checkpoint before exiting — the
thing users are told to do on preemptible capacity — holds its pod object longer and therefore
leaks more quota. A task that ignores shutdown entirely and gets SIGKILLed is affected least.
A trainer that checkpoints for 60 s leaks ~700 charges.
Proposed fixes, in order of what they buy
- Treat a non-terminal pod with a
deletionTimestamp as "still shutting down", not as a
failure (plugin_manager.go:219). A graceful delete is a normal, time-bounded state. Keep
observing until the object is gone — the NotFound path already produces
ResourceDeletedExternally (:164) — and only then decide what to do. Alternatively, restore
the v1 semantics and start a new attempt (new name) on a SYSTEM failure instead of
re-launching under the same name. Either one removes the cause: no Create is ever issued
against a name that is known to be occupied, so nothing leaks. This is the fix that matters.
- Do not classify an admission/quota
403 as a user error, and add backoff to retryable
failures (plugin_manager.go:119, taskaction_controller.go:832-847). A ResourceQuota
rejection is a transient platform condition; v1's WaitingForResources + backoff is the prior
art. Even with fix 1, an unrelated quota exhaustion still burns a task's whole budget in
milliseconds today.
launchResource should not swallow AlreadyExists into Queued (:118). It reports
"submitted" for a pod it did not submit. Note this is hardening, not a fix on its own: the leak
is per failed Create, and the recovery would still issue one per lap.
- Do not reset
Status.SystemFailures on a transition produced by the same recovery cycle
(taskaction_controller.go:827). Also hardening only — and not safe alone: bounding the
loop at 3 without fix 1 turns every graceful deletion into a permanent
MaxSystemFailuresExceeded failure within a second.
Happy to open a PR for (1) if the maintainers agree on the direction (observe-until-gone vs.
new-attempt).
Related: #8067 - with the DisruptionTarget condition present (set by KAI before the
delete, so it is on the pod for the whole grace window), the executor can also tell "preempted,
shutting down" from "deleted behind my back" in this very branch.
Summary
When a task pod is deleted gracefully while its container does not exit immediately, the
TaskActionreconciler enters a livelock: it alternates between reportingUnexpectedObjectDeletionand re-creating a pod under the same name that already exists, atroughly 12 iterations per second, for as long as the pod object survives.
MaxSystemFailurescannot bound it, because a counter reset in the loop's other half clears it every lap.
Each iteration issues a
Createthat fails withAlreadyExists, and each of those failed createsleaks one pod's worth of
ResourceQuotausage (a known Kubernetes behaviour — admission chargesbefore the registry create and does not roll back: kubernetes/kubernetes#70563,
kubernetes/kubernetes#106842, kubernetes/kubernetes#110080). In a namespace with a
podsquota,the loop therefore exhausts the quota in seconds. The task's own next attempt is then rejected with
403 Forbidden, which the executor classifies as a user error and retries with no backoff,burning every remaining attempt in a few hundred milliseconds without ever creating a pod.
Net effect: a routine graceful pod deletion kills the task and temporarily exhausts a quota shared
by every other workload in the namespace.
This is a regression against v1 propeller, which has the same
UnexpectedObjectDeletionbranch butrecovers differently (see "How v1 handled the same situation").
Versions
Observed on v2.0.35 (
cr.flyte.org/flyteorg/flyte-binary-v2:v2.0.35), clientflyte2.5.16.Every code path cited below is present in v2.0.49; line numbers are from that tag.
Reproduction
A
podsResourceQuota in the task namespace (any modest cap):A task that does not exit promptly on
SIGTERM, withretries=5so the budget burn is visible:Then, once the pod is
Running:A plain
kubectl delete podis enough - it is an ordinary graceful delete, the same call ascheduler makes when it preempts. We hit this with NVIDIA KAI; kube-scheduler preemption,
cluster-autoscaler scale-down and node drain all take the same path.
The explicit
SIG_IGNis for determinism, not necessity: in our runs a task with no signalhandling at all also held its pod
Runningwith adeletionTimestampfor the full 30 s graceperiod, so the default case is already affected. Ignoring the signal just removes any dependency on
how the runtime happens to handle it.
Observed: ~90
UnexpectedObjectDeletionerrors in 8s,used.podsclimbing to 100 in lockstep(one increment per ~80 ms), then attempts 2–6 in 300 ms, then the action
FAILED. RaisingterminationGracePeriodSecondsscales the leak linearly.Mechanism
All line numbers are v2.0.49.
1. A graceful delete is read as an unexpected one.
executor/pkg/plugin/k8s/plugin_manager.go:219A pod under a graceful delete has a
deletionTimestampand remains non-terminal for its wholegrace period, so this branch is true continuously for up to
terminationGracePeriodSeconds— it is not a one-shot condition.2. The recovery re-creates a pod under the same name, and misreports success.
executor/pkg/controller/taskaction_controller.go:812-814routes the system-retryable failure toresetPluginResource(:501), which aborts and clearsStatus.PluginStatebut keepsStatus.Attempts. The next reconcile therefore takes thePluginPhaseNotStartedpath intolaunchResourcewith the same generated name. There,plugin_manager.go:117-143:The still-terminating pod makes
CreatereturnAlreadyExists, which is swallowed, and thetransition reports
Queuedas though a fresh pod had been submitted.The
AbortinsideresetPluginResourcedoes not help: it issues a plainDeletewith no graceoverride (
plugin_manager.go:465), and a second graceful delete of an already-terminating pod is ano-op.
3.
MaxSystemFailurescan never bound the loop.taskaction_controller.go:827This sits after the early return at
:812, so the loop alternatesSystemFailures++(:565) →Queued→ reset →SystemFailures++→ … and never reachesDefaultMaxSystemFailures = 3(:69).The loop is self-triggering: the controller is registered with
For(&TaskAction{})and nopredicate (
:1420-1423), and both halves of the loop writeTaskActionstatus(
recordSystemErrorpersists the counter, theQueuedhalf persistsPluginState). Each write isa watch event that re-enqueues the same object, so the cadence is API latency, not the requeue
duration (which is why the observed ~12 Hz is far above the 10 s default requeue, and why making
the requeue configurable in #7770 does not affect this).
4. The resulting quota rejection is charged to the user, with no backoff.
plugin_manager.go:119-121A ResourceQuota rejection is
403 Forbidden, so it becomes a USER-kind retryable failure. Thein-place restart path (
taskaction_controller.go:832-847) incrementsStatus.Attemptsandre-enters immediately — there is no backoff on retryable failures anywhere in this path — so the
whole budget is spent at API speed. The action's final
error_infois:No pod was created for attempts 2–6; there is no pod object or event for any of them.
How v1 handled the same situation
v1 propeller has the identical deletion branch
(
flytepropeller/pkg/controller/nodes/task/k8s/plugin_manager.go:346-352onmaster), with acomment naming the intended cases: a vanished kubelet leaving pods stuck behind the flyte
finalizer, and "when a user deletes a Pod directly". Two things made it harmless there:
nodes/executor.go:830,isEligibleForRetry), so the next pod had a new name andAlreadyExistscould never collidewith the dying one. v2's
resetPluginResourcereuses the attempt number, which is the root ofthe livelock.
exceeded quotaonCreatewas recognised explicitly and mapped toWaitingForResourcesbehind a backoff controller (
plugin_manager.go:247-258onmaster,backoff.IsResourceQuotaExceeded). v2 dropped that and lets the genericIsForbiddenbranchturn it into a user error.
Why this matters beyond one cluster
drain, cluster-autoscaler scale-down, a manual
kubectl delete pod.workload's pod creation in that namespace is rejected.
SIGTERMto checkpoint before exiting — thething users are told to do on preemptible capacity — holds its pod object longer and therefore
leaks more quota. A task that ignores shutdown entirely and gets
SIGKILLed is affected least.A trainer that checkpoints for 60 s leaks ~700 charges.
Proposed fixes, in order of what they buy
deletionTimestampas "still shutting down", not as afailure (
plugin_manager.go:219). A graceful delete is a normal, time-bounded state. Keepobserving until the object is gone — the
NotFoundpath already producesResourceDeletedExternally(:164) — and only then decide what to do. Alternatively, restorethe v1 semantics and start a new attempt (new name) on a SYSTEM failure instead of
re-launching under the same name. Either one removes the cause: no
Createis ever issuedagainst a name that is known to be occupied, so nothing leaks. This is the fix that matters.
403as a user error, and add backoff to retryablefailures (
plugin_manager.go:119,taskaction_controller.go:832-847). A ResourceQuotarejection is a transient platform condition; v1's
WaitingForResources+ backoff is the priorart. Even with fix 1, an unrelated quota exhaustion still burns a task's whole budget in
milliseconds today.
launchResourceshould not swallowAlreadyExistsintoQueued(:118). It reports"submitted" for a pod it did not submit. Note this is hardening, not a fix on its own: the leak
is per failed
Create, and the recovery would still issue one per lap.Status.SystemFailureson a transition produced by the same recovery cycle(
taskaction_controller.go:827). Also hardening only — and not safe alone: bounding theloop at 3 without fix 1 turns every graceful deletion into a permanent
MaxSystemFailuresExceededfailure within a second.Happy to open a PR for (1) if the maintainers agree on the direction (observe-until-gone vs.
new-attempt).
Related: #8067 - with the
DisruptionTargetcondition present (set by KAI before thedelete, so it is on the pod for the whole grace window), the executor can also tell "preempted,
shutting down" from "deleted behind my back" in this very branch.