Skip to content

Wait out connection blips and recover as soon as the network returns - #385

Merged
MaggieAppleton merged 10 commits into
mainfrom
design/connection-resilience
Oct 8, 2026
Merged

MaggieAppleton merged 10 commits into
mainfrom
design/connection-resilience

Conversation

@MaggieAppleton

@MaggieAppleton MaggieAppleton commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Why

A half-second socket blip set off three alarms at once: Chat inserted a "Connection lost · Reconnect" row that pushed the transcript up (then "Synchronizing… · Retry"), the document showed its Reconnecting pill, and decision-card actions dimmed. Then it all snapped back. During a real outage the Chat composer was locked, so you could not even draft, and recovery waited for the backoff timer (up to about 22 s) after the network was already back.

What changed

  • Grace period. New useConnectionNotice in @chopin/editor: nothing shows until a loss outlasts 1.5 s, then reconnecting, then offline 5 s later. The document lock and status pill, decision-card dimming, and Chat all use it. Reopening the document and sending still read the real connection. Edits typed during the grace period wait in the provider's outbox and are replayed on resume.
  • Chat. The connection notice row is gone. A text-xs status sits in the composer footer: "Reconnecting…", then "Offline" with a ghost "Reconnect" button. It fades in with --duration-fast, and there is no fade under reduced motion. Nothing moves. "Synchronizing…/Retry" is dropped.
  • Draft while offline. The composer and mode switch stay editable. Only Send is disabled, and submit also checks the live socket, so typed text is never lost or silently dropped.
  • Fast recovery. Wire skips the rest of its backoff on online, window focus, and visibilitychange to visible. Only a wire waiting on its retry timer reacts, so an attempt already in flight is left alone. The listeners are removed on dispose or deletion.
  • Locked editor. No offline editing. The read-only lock and the ink dimming from Show document connection status in the document header #297 (data-plan-offline) now start only after the grace period ends, so neither appears during a blip.

packages/editor/src/status.tsx changes minimally: an optional onReconnect prop. When it is given, the offline alert offers Reconnect instead of Reload. The header pill from #297 receives the connection only after the grace period, so it reads Reconnecting… then Offline, matching Chat.

After review

  • Epoch rotated during a blip: PlanProvider#open used to merge the new epoch's state into the old document. It now rebuilds from the server, the same way plan:reset does. If unsent edits were dropped, the document header says "Edits not saved", with the full explanation ("Your last edits couldn't be saved because the document changed while you were offline.") in its tooltip and accessible text, and a Dismiss button. It sits in the status area, so it covers no decision card. This also closes the few-millisecond window that main had.
  • Wake: focus and visibility events wake the wire at most once per second after the last attempt, and they keep the backoff. Only online resets it.
  • No flash past the grace period: one retry is always made 1.2 s after a loss, so a short outage reconnects before the 1.5 s grace runs out. The lock and the status still change together.
  • Decision actions: Save, Discard and Cancel say "Not connected. Try again when you're back online." when the request never reached the server.
  • Chat Send: the button no longer dims during a blip. Enter waits for the connection, for at most the grace period, and then sends.
  • One offline state: Chat uses the header's words and tone (a warning dot for "Reconnecting…", red "Offline"). The header's offline action is now Reconnect, which reconnects in place. Reload appears only after three failed reconnects in one outage, and asking to reconnect restarts the backoff. On a phone, where the header is out of view, the Chat footer offers Reconnect too. Chat stores an unsent message in sessionStorage as it is typed, so no reload loses it. Stop and Resume Chopin no longer dim during a blip; a press waits and is sent once the connection is back.
  • connectionShown is renamed to treatAsConnected.

Lost edits in the header

Chat and header now share one offline language:

Outage after review

Screenshots

Blip, before (main): the notice row appears and the pill shows
Blip before

Blip, after: nothing changes
Blip after

Long outage, before: the composer is locked and shows "Connection lost"
Outage before

Long outage, after: the draft is kept, the footer says Reconnecting… then Offline, and the document ink is dimmed
Outage after

Phone, before / after
Phone before
Phone after

Testing

  • bun test apps/web/src/wire.test.ts: 14 pass. New: the backoff is skipped on online, and a wake while connected is ignored.
  • bun test --timeout 60000 packages/editor apps/web: 1030 pass, 0 fail (after rebasing).
  • bun run types passes. bun scripts/check-design-record.ts passes.
  • New e2e/connection.e2e.ts (routeWebSocket): a short blip shows nothing; a long outage shows one calm state; Chat text survives and sends; the online event reconnects immediately. e2e/editing.e2e.ts is updated so the lock test holds the socket down past the grace period. Local E2E is suspended for this run, so CI runs these.
  • Manual check against a fake-GitHub server: after the online event the editor unlocked in about 15 ms. The 8795 reference took 0.7–6 s.

Final integration

Rebased onto #379 main with the latest authored commits preserved. bun run ci, bun run types, 224 focused Chat/provider tests, and the production build pass (247,156 B raw / 77,951 B gzip initial JavaScript). The editor design-contract review hash is renewed. Fresh hosted validation, browser E2E, and container checks all passed on final head 3fa30fe3.

🤖 Generated with Claude Code

@MaggieAppleton
MaggieAppleton force-pushed the design/connection-resilience branch from 074bc22 to 06b8042 Compare October 8, 2026 02:09
@MaggieAppleton

Copy link
Copy Markdown
Collaborator Author

Hold — please don't merge yet. Changes connection/lock timing and edit replay on reconnect; a sync-focused review is in progress. I'll post the verdict here.

🤖 Generated with Claude Code

@MaggieAppleton

Copy link
Copy Markdown
Collaborator Author

Review: Changes needed

The visual and Chat side is good. A blip under 1.5 s shows nothing, the footer status is calm, and the transcript does not move on desktop or phone. The draft survives a long outage, and the label fades in cleanly with opacity only. A 1 s blip replays its edits to the peer exactly once. Retrying a Chat send whose acknowledgement was lost reuses the requestId, so it does not duplicate. One sync bug blocks this, and the wake logic needs one fix.

I tested with my own headless Playwright: two users, routeWebSocket proxy, the PR server on 8898.

Blocking

  1. An epoch change during a blip silently forks the document. packages/editor/src/provider.ts #open (the rotated branch). Ana goes offline and types " STALE" during the grace period, which is now allowed. Meanwhile the server rebuilds after Ben sends an invalid batch, so the epoch rotates. On reconnect, #open applies the new epoch's state onto Ana's old doc and clears the outbox. That avoids a stale replay. But "STALE" stays visible in Ana's editor, and the pill says "Reconnected". Ben and a fresh load never get it. Worse, everything Ana types afterwards is sent with the new epoch and acknowledged, but never applied, because it depends on the missing items. In my run " AFTER" got 6 acks and Ben saw 0. Ana is silently diverged until she reloads.
    On main the editor locks the moment the socket drops, so this window was only a few milliseconds of unacknowledged updates. The grace period turns it into ordinary typing. This is the AGENTS.md trap "replay unacknowledged updates only when the epoch is still compatible": not replaying is right, but the local state has to go too.
    Fix: on a rotated open, take the #reset path instead of merging into the old doc. Call onReset so the editor rebuilds from the server state. If the outbox was not empty, tell the person their last few edits could not be applied. Add a provider unit test (outbox not empty + open reply with a new epoch → onReset, nothing replayed) and a connection.e2e.ts case.

Should fix

  1. #wake resets the backoff on every focus or visibility event. apps/web/src/wire.ts #wake sets #attempts = 0 and connects immediately. While the server is down, a focus/visibility storm (every 100 ms for 6 s) caused 62 socket attempts and 63 HTTP probes in 20 s. The baseline is 5 attempts with gaps of 459, 583, 1430, 4023 and 8742 ms. After the storm the backoff starts again at about 250 ms. During a deploy outage, every tab switch on every client starts a fresh fast-retry cycle. Suggest: don't reset #attempts on a wake (or reset it only on online), and ignore a wake within about BASE_DELAY of the last attempt. Connected and in-flight wakes are correctly ignored (0 attempts).

Minor

  1. The status flashes just past the grace period. A 1 s outage usually reconnects at about 1.6–1.7 s, because the second retry lands 0.5–1.5 s after the first failure. I measured "Reconnecting…", the notice pill and an editor lock for about 100 ms, then all of it disappeared. Keystrokes in that 100 ms are dropped. Consider a short minimum display or hysteresis, or wake once at the grace boundary so the retry lands before it.
  2. Save on a decision card during the grace period fails loudly. Cards still get connected: true (room-workspace.tsx connectionShown), so Save is enabled. Clicking it 300 ms into a blip shows a red "Couldn't save / Could not submit these answers." alert. That is the alarm this PR set out to remove, and the copy does not mention the connection. The selection survives and Save works after reconnect. Better: wait briefly for the reconnect before failing, or at least say "Not connected. Try again when you're back online."
  3. Enter in Chat during the grace period is silently ignored. Send dims immediately because it reads the live socket, so a blip still changes the composer. Fine if intended, but it contradicts "nothing changes".
  4. Two offline indicators disagree. When offline, the header says Offline (red) with Reload, and the Chat footer says Offline (orange dot) with Reconnect. That is one state with two tones and two different actions. Consider one action, or matching tone.
  5. Naming. connectionShown is true when nothing is shown, which reads backwards. Something like treatAsConnected would be clearer.

Checks

gh pr checks 385: e2e pass, container pass. format, lint, types, tests fails only on the listed reviewed dynamic owner packages/editor/src/plan-editor.tsx (dynamic-editor.json). That is acceptable per the brief. Local bun test for wire, connection-notice and provider: 27 pass.

@MaggieAppleton
MaggieAppleton force-pushed the design/connection-resilience branch from a7f0e18 to bb1de85 Compare October 8, 2026 03:29
@MaggieAppleton

Copy link
Copy Markdown
Collaborator Author

Thanks. All seven points are addressed, and the branch is rebased onto main.

  1. Epoch fork (blocker), fixed.
    • Fix: in PlanProvider#open, a reply with a different epoch now goes down the plan:reset path (#discard). The client clears the outbox, replays nothing and calls onReset, so the editor rebuilds from the server's state instead of merging into the old document. This also closes the few-millisecond window that main had.
    • Notice: onReset(reason, lost) reports whether unacknowledged edits were dropped. If they were, the editor shows "Your last edits couldn't be saved because the document changed while you were offline." with a Dismiss button. It sits over the prose, so it moves nothing.
    • Tests:
      • provider.test.ts: a non-empty outbox plus a rotated reply calls onReset("replaced", true) and sends only plan:open; an empty outbox reports lost: false.
      • connection.e2e.ts: Ben sends an invalid batch while Ana is offline and typing. Ana's stale text is gone, the notice shows, and what she types afterwards reaches Ben.
    • Manual check: two sessions on the fake server; the stale text reached nobody, and later edits reached the other session.
  2. Wake storm, fixed. Focus and visibility no longer reset #attempts, and they wake the wire at most once per second after the last attempt. online still reconnects at once and resets the backoff. New wire.test.ts case: 30 focus/visibility events over 3 s cause at most 4 attempts.
  3. Flash past the grace period, fixed. One retry is always made 1.2 s after a loss (RESCUE_AT), so a short outage reconnects before the 1.5 s grace runs out. The lock and the status still change together. Test: with the jitter forced to its maximum, the wire reconnects within 1.5 s.
  4. Save during the grace period. When Save, Discard or Cancel fail because the request never reached the server, they now say "Not connected. Try again when you're back online." instead of "Could not submit these answers." Unit-tested.
  5. Send during the grace period. The button no longer dims. Enter waits for the connection (and fresh history), for at most the grace period, then sends. After the grace period, Send is disabled and the draft stays. Covered by connection.e2e.ts and by the updated chat-references and smoke tests.
  6. One offline language. Chat now uses the header's words and tone: a warning dot for "Reconnecting…" and red "Offline". Chat's own Reconnect button is gone, so the header's Reload is the only action.
  7. Naming. connectionShown is now treatAsConnected.

CI: e2e and container pass. format, lint, types, tests fails only on the reviewed dynamic owner plan-editor.tsx; the new hash is in the PR body. Local: bun test packages/editor apps/web packages/question gives 1185 pass.

@MaggieAppleton

Copy link
Copy Markdown
Collaborator Author

Review: Looks good

I re-checked at 19e57220 with my own headless Playwright: two users, a routeWebSocket proxy, and the PR server on 8898. Every earlier finding is fixed.

  • Epoch fork: fixed. Ana typed " STALE" while offline, and Ben forced a rebuild with an invalid batch. "STALE" reached nobody: it is not in Ana's view, Ben's view or a fresh load. The notice "Your last edits couldn't be saved because the document changed while you were offline." appears. Ana's later " AFTER" reached Ben and the server.
  • Focus storm with the server down: fixed. Events every 100 ms for 6 s now cause 7 attempts and 8 probes in 20 s, against 62 before. Wakes are capped at about 1 per second, and backoff resumes afterwards (the next gap is 13.5 s). Wakes while connected cause no attempts.
  • 1 s outage: no flash. The rescue attempt lands at 1.2 s and the connection is back at about 1.25 s. The lock, pill and footer never show. " BLIP" reaches Ben exactly once.
  • Chat during a blip. Enter during the grace period waits, then sends once. Retrying a send whose acknowledgement was lost reuses the requestId, so there is no duplicate.
  • Offline for more than 5 s. The header and Chat use the same language and tone. The header's Reconnect brings the connection back in about 25 ms, and the draft is kept. The transcript does not move on desktop or phone.
  • Decision Save during the grace period now says "Not connected. Try again when you're back online." The selection is kept, and Save works after reconnecting.

Non-blocking nits:

  • On a phone, in the Chat tab, the footer says "Offline", but the Reconnect action only exists in the Document tab's header. Automatic retries and online/focus wakes still recover, so this is not a blocker.
  • The lost-edits notice sits over the bottom right of the document and can cover a decision card's Save button until it is dismissed.

Checks: bun run ci locally fails only on the reviewed dynamic owner plan-editor.tsx, which is acceptable per the brief. The touched unit suites pass (147 tests).

@MaggieAppleton

Copy link
Copy Markdown
Collaborator Author

Hold lifted.

@MaggieAppleton

Copy link
Copy Markdown
Collaborator Author

CI follow-up: after my review, the e2e job finished red on the merge with main (run 37725777679). The hold is lifted on design and sync grounds, but this needs fixing before the PR can merge.

  • e2e/chat-mentions.e2e.ts:211, "without a Planner a manually addressed message gets a local notice". This fails on every retry. After Send, the message @chopin are you there? never appears. The test comes from 70c71e0 ("Say so when the Planner is unavailable"), which landed on main after this branch's base and also changes apps/web/src/chat/chat.tsx. The new Send-waits-during-grace path probably doesn't combine cleanly with it. Please rebase onto main and reconcile the two.
  • e2e/editing.e2e.ts:215, "a lost connection is said in the document header and the composer". This was flaky: ready() timed out once, with the editor still locked after reconnecting. Worth checking that it is stable after the rebase.

@MaggieAppleton
MaggieAppleton force-pushed the design/connection-resilience branch 2 times, most recently from ed73942 to 9221187 Compare October 8, 2026 04:36
@MaggieAppleton

Copy link
Copy Markdown
Collaborator Author

Rebase re-verified: Looks good. At 9221187 the merged provider sends a rotated-epoch open through #discard, which clears the outbox and all resend/pacing state via #stopTimers, then calls onReset(reason, lost). A same-epoch reopen still replays through #replay with the send counts reset. I reran my epoch-fork repro: the stale text reached nobody, the 'Edits not saved' notice showed, and later edits reached the peer and the server. A 1 s blip, run twice, reconnected at about 1.2 s with no lock or notice, and the text arrived exactly once. Reconnect resetting the backoff only happens on a click; the header caps it at RECONNECTS_BEFORE_RELOAD and focus/visibility wakes stay limited, so it can't cause a storm. provider and wire unit tests pass.

MaggieAppleton and others added 10 commits October 8, 2026 06:04
A socket drop now shows nothing for 1.5 s, then one calm status in the
Chat composer footer (Reconnecting… then Offline) instead of a notice row
that pushed the transcript. The wire reconnects at once on online, focus
and visibility, the Chat draft stays editable while offline, and the
locked document dims its ink.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… composer

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A document rebuilt while a client was away is now rebuilt locally too,
rather than merged into, so edits made during a blip cannot fork it; the
person is told when unsent edits were dropped. Focus and visibility wake
the wire at most once a second and keep the backoff, and one attempt lands
inside the grace period. A send during a blip waits for the connection,
decision actions say when they failed for want of one, and Chat's
offline status matches the document header's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ss reloads

The header's offline action now reconnects in place. Reload is offered
only after three reconnects fail in one outage, and Chat keeps an unsent
message in sessionStorage across any reload. Stop and Resume Chopin stay
usable during a blip and are sent once it is over.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… restart backoff on request

The notice for edits a rebuild dropped now sits with the document status
in the header, where it covers no decision card. Chat offers Reconnect
when the header is out of view. Asking to reconnect starts the backoff
over, so a few failed asks never leave the next automatic attempt waiting
out the cap. The no-Planner mention test finds the sent message by its
raw text, since the transcript drops the leading mention.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaggieAppleton
MaggieAppleton force-pushed the design/connection-resilience branch from 9221187 to 3fa30fe Compare October 8, 2026 05:05
@MaggieAppleton
MaggieAppleton merged commit 1c93928 into main Oct 8, 2026
3 checks passed
@MaggieAppleton
MaggieAppleton deleted the design/connection-resilience branch October 8, 2026 05:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant