Summary
When a Connected Services pool rotates because another session hit a usage limit, an idle Codex session adopts the new member in its metadata and credential file but keeps serving on the previous account, so its next turn fails and the pool selector parks it with no_eligible_member while the new member is healthy.
What happened (current behavior)
The idle session's binding advanced to the new pool member and its CODEX_HOME/auth.json was rewritten to that member's account, but its running Codex app-server received no account/login/start, so the next turn was answered by the previous account and rejected with usage_limit_exceeded, and the failure was then reported under the new member's profile id.
Expected behavior
A session that adopts a new auth generation serves its next turn on the newly selected member, and a usage-limit failure is attributed to the account that actually served the request.
Reproduction steps
- On a self-hosted relay, put at least two Codex profiles from different ChatGPT accounts into one Connected Services pool, with distinguishable plans so the serving account is identifiable from the provider
rate_limits payload.
- Start two Codex sessions on that pool in
appServer backend mode. Call them A and B.
- Run session A until a turn completes, then leave it idle.
- Run session B until the provider reports a usage limit, so the daemon rotates the group to the next member and advances the group generation.
- Confirm session A now reports the new member in its metadata and that its
CODEX_HOME/auth.json holds the new member's account.
- Send a message to session A. The turn fails with a usage limit whose reset time and plan type belong to the previous account, and the daemon logs the switch attempt as
"resultStatus":"no_eligible_member" with the new member excluded as current_active.
Severity
medium
Frequency
once
Happier version
0.2.11-dev.1
Platform
Linux x64 (Debian 13 container), Codex CLI 0.147.0, appServer backend mode
Server version
0.2.11-dev.1
Deployment type
self-hosted
What changed recently?
No response
Diagnostics ID
No response
Additional context
Observed timeline, working control, supporting evidence and remaining uncertainty
T0 is the moment the affected session's group binding advanced to the new member. The session had ended its previous turn about 7 minutes before T0 and was idle. Its Codex app-server process started about 55 minutes before T0 and was never restarted.
Observed:
- T0: the affected session's group binding advances to the new member and the group generation increments. The rotation was triggered by the other session's usage limit, not by this one.
- T0+14s: its
CODEX_HOME/auth.json is rewritten. The credential's own account id and its chatgpt_plan_type claim both belong to the new member, so the file on disk is correct.
- Around the same second, Codex's app-server request trace for that process shows only
account/read and account/rateLimits/read. No account/login/start is received before, during or after the swap.
- T0+2m06s: the next turn is rejected with
usage_limit_exceeded. The provider payload carries the previous account's plan_type and a 5 hour window reset. The new member has no 5 hour window at all and its weekly window resets four days later, so the response could not have come from it.
- T0+2m08s: the runner reports the failure with the new member's profile id. The selector excludes that member as
current_active on fresh quota evidence showing it not exhausted, marks every other member quota_exhausted, and returns no_eligible_member. The recovery record is armed as waiting with a wake time about 3h48m later, taken from an unrelated member's reset.
Working control, same pool, same generation, same host, same CLI version and same Codex binary: the session that hit the limit itself adopted the new member inside its existing app-server process without a restart, and its next turn 6 seconds later carried the new member's plan type and weekly window. The credential apply mechanism therefore works; the divergence is specific to a session that adopts a generation it did not trigger.
Supporting evidence for the same boundary: at T0+2s, 12 seconds before the credential file was written and 2 minutes before the turn that the previous account actually served, the runtime identity probe for the affected session already reported the new member's provider account id with "probeStatus":"verified" and "proofKind":"runtime_exact". The affected Connected Services corridor in the observed build is identical to public 0.2.11-dev.1 at 8cf3c40, where that probe returns the identity cached from the applied credential (source: 'applied_credential') and only refreshes it from the live account when the cached value does not match the refresh selection, so an identity that matches intent but not the live session is returned unverified against the provider.
Remaining uncertainty: whether the credential apply on the idle session reported success or failure is not recoverable, because the RPC result payload is not logged on either side and the daemon records no switch attempt for a session that only observes a new generation. Codex's own log store had already pruned the control session's window, so the presence of account/login/start in the working path was not verified in the same artifact.
Workaround: sending another message after the failure runs on the new member, because the failed turn resets the Codex session first. Until then the session stays parked.
Related: #276 and #277 also end with a session that never resumes, but both stop at the pool policy phase with auto_switch_disabled and a recovery record that carries no wake time. Here the policy allows the switch, the record does carry a wake time, and the session parks because the failure is attributed to a member that did not serve the request.
Summary
When a Connected Services pool rotates because another session hit a usage limit, an idle Codex session adopts the new member in its metadata and credential file but keeps serving on the previous account, so its next turn fails and the pool selector parks it with
no_eligible_memberwhile the new member is healthy.What happened (current behavior)
The idle session's binding advanced to the new pool member and its
CODEX_HOME/auth.jsonwas rewritten to that member's account, but its running Codex app-server received noaccount/login/start, so the next turn was answered by the previous account and rejected withusage_limit_exceeded, and the failure was then reported under the new member's profile id.Expected behavior
A session that adopts a new auth generation serves its next turn on the newly selected member, and a usage-limit failure is attributed to the account that actually served the request.
Reproduction steps
rate_limitspayload.appServerbackend mode. Call them A and B.CODEX_HOME/auth.jsonholds the new member's account."resultStatus":"no_eligible_member"with the new member excluded ascurrent_active.Severity
medium
Frequency
once
Happier version
0.2.11-dev.1
Platform
Linux x64 (Debian 13 container), Codex CLI 0.147.0,
appServerbackend modeServer version
0.2.11-dev.1
Deployment type
self-hosted
What changed recently?
No response
Diagnostics ID
No response
Additional context
Observed timeline, working control, supporting evidence and remaining uncertainty
T0 is the moment the affected session's group binding advanced to the new member. The session had ended its previous turn about 7 minutes before T0 and was idle. Its Codex app-server process started about 55 minutes before T0 and was never restarted.
Observed:
CODEX_HOME/auth.jsonis rewritten. The credential's own account id and itschatgpt_plan_typeclaim both belong to the new member, so the file on disk is correct.account/readandaccount/rateLimits/read. Noaccount/login/startis received before, during or after the swap.usage_limit_exceeded. The provider payload carries the previous account'splan_typeand a 5 hour window reset. The new member has no 5 hour window at all and its weekly window resets four days later, so the response could not have come from it.current_activeon fresh quota evidence showing it not exhausted, marks every other memberquota_exhausted, and returnsno_eligible_member. The recovery record is armed aswaitingwith a wake time about 3h48m later, taken from an unrelated member's reset.Working control, same pool, same generation, same host, same CLI version and same Codex binary: the session that hit the limit itself adopted the new member inside its existing app-server process without a restart, and its next turn 6 seconds later carried the new member's plan type and weekly window. The credential apply mechanism therefore works; the divergence is specific to a session that adopts a generation it did not trigger.
Supporting evidence for the same boundary: at T0+2s, 12 seconds before the credential file was written and 2 minutes before the turn that the previous account actually served, the runtime identity probe for the affected session already reported the new member's provider account id with
"probeStatus":"verified"and"proofKind":"runtime_exact". The affected Connected Services corridor in the observed build is identical to public0.2.11-dev.1at 8cf3c40, where that probe returns the identity cached from the applied credential (source: 'applied_credential') and only refreshes it from the live account when the cached value does not match the refresh selection, so an identity that matches intent but not the live session is returned unverified against the provider.Remaining uncertainty: whether the credential apply on the idle session reported success or failure is not recoverable, because the RPC result payload is not logged on either side and the daemon records no switch attempt for a session that only observes a new generation. Codex's own log store had already pruned the control session's window, so the presence of
account/login/startin the working path was not verified in the same artifact.Workaround: sending another message after the failure runs on the new member, because the failed turn resets the Codex session first. Until then the session stays parked.
Related: #276 and #277 also end with a session that never resumes, but both stop at the pool policy phase with
auto_switch_disabledand a recovery record that carries no wake time. Here the policy allows the switch, the record does carry a wake time, and the session parks because the failure is attributed to a member that did not serve the request.