Skip to content

Connected Services: capacity recovery reports success while the interrupted Codex session stays idle #344

Description

@karolzlot

Summary

A Codex provider-capacity failure arms automatic recovery, but Happier reports the session as resumed while it remains idle and requires a new user message.

What happened (current behavior)

After server_overloaded without reset or retry-after metadata, the daemon armed temporary recovery, logged a successful resume about one second later, cancelled the recovery intent on its first attempt, and left the live runner waiting with an empty input queue.

Expected behavior

Automatic capacity recovery should either start one safe continuation after backoff or remain visibly pending or action-required instead of reporting success without new provider activity.

Reproduction steps

  1. Run a Codex turn through a Connected Services profile or pool.
  2. Let the provider terminate the turn with server_overloaded and no reset or retry-after metadata.
  3. Observe temporary_retry_armed, followed by Temporary throttle recovery resumed session after the default backoff.
  4. Observe that no successor turn starts, the runner waits with no pending input, and the recovery intent is already cancelled.

Severity

medium

Frequency

always

Happier version

CLI and daemon 0.2.10-dev.83

Platform

Linux x64 (Debian 13)

Server version

0.2.10-dev.76

Deployment type

self-hosted

What changed recently?

No response

Diagnostics ID

No response

Additional context

Verified recovery boundary and related issue

At public source cce6894, provider capacity is routed to temporary retry. Without provider timing metadata, the scheduler retries after one second, verifies only that the tracked session still exists, and calls the existing-session spawn path. That path adopts the already-live runner and returns success after nudging its empty pending queue, so the scheduler cancels the intent without a successor turn or provider-activity proof.

Related: #309 covered provider errors that bypassed capacity classification. Here server_overloaded is classified correctly and reaches the temporary-retry owner; the failure occurs afterward when runner adoption is treated as successful continuation.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs:reporterWaiting for new information or confirmation from an external participant.priority:p1High prioritystage:stabletype: bugBug or regression

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions