Skip to content

Add child working directories and durable parent loop wakeups - #1211

Merged
SamSaffron merged 2 commits into
SamSaffron:mainfrom
sam-saffron-jarvis:feat/subagent-migration-gaps
Oct 6, 2026
Merged

SamSaffron merged 2 commits into
SamSaffron:mainfrom
sam-saffron-jarvis:feat/subagent-migration-gaps

Conversation

@sam-saffron-jarvis

@sam-saffron-jarvis sam-saffron-jarvis commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Scope

Continue the native spawn_agent migration gaps on base a2737d6b; this does not route native execution through jobs-v2.

  • Optional cwd: inherit the parent's actual directory by default; resolve relative paths there; require an existing directory; persist the resolved path for continuation without widening workspace/approval authority.
  • Optional notify_when_done: reactivate the original parent agent loop, not emit a raw child report. Delivery is currently supported only for persisted web parents; Telegram, terminal chat, ask, and sessionless hosts explicitly reject the option.
  • Durable generation-scoped completion/collection and retained media; serialize/coalesce parent continuations behind normal response admission.
  • Surface host restart/reload interruption without automatically restarting child work. Host lifecycle turns cannot call continue_agent; recovery requires a subsequent user-authorized turn.
  • No execution deadline: the newest user ruling supersedes the earlier execution-timeout request. wait/deprecated timeout and max_wait (0–3600 seconds) bound caller collection only. The configured turn limit (default 500) remains the execution budget. Native collection pauses parent inactivity; a parent inactivity timeout/lease loss does not terminate independently owned detached children.

No merge, deployment, live session database access, or Jarvis configuration change.

Authorized reviews and fixes

Exactly the two already-running read-only reviews were collected, with no additional reviews or paid-model calls:

  • Astra: chatgpt:gpt-6-astra-medium, run run_0gR3wk4JpMDJ — completed, changes requested; four substantive findings.
  • Opus: claude-bin:opus-max, run run_A6cqN36FYSFl — completed, changes requested; blocking and important lifecycle/provider findings.

Both reviewed 376dadbdc432816b663b5bccea3d149025611337. The follow-up commit is ec534ba094821625349dea4dd870be3fe01fa2a2; the reviewers have not re-reviewed or approved that follow-up.

Addressed with targeted regressions:

  1. Build host wakes through the normal parent request projection: authorized tool specifications, max turns, search, parallel tools, tool map, persisted model/reasoning settings.
  2. Use a durable marked synthetic provider-visible user turn rather than a developer-only turn. Claude-bin fresh/resumed, Gemini, and grok-bin serializer tests retain the event. Child output/error is not promoted to developer authority: the event contains lifecycle metadata only; results/media arrive through wait_agent.
  3. Revalidate event generations/collection/stop under local session admission and again in the durable lease-admission transaction. Retry when a concurrent completion changes the wake signal; retry restart events when expired response ownership is recovered.
  4. Preserve acknowledgment ownership and internal-wake metadata across reload. Restore interrupted user responses before reconciling pending wakes.
  5. Distinguish graceful host shutdown from explicit user cancellation. Explicit stop suppresses active session children and completions from the stopped turn, plus generations offered by a stopped host wake, without rewriting earlier completed lifecycle reasons.
  6. Never recreate deleted/archived parent sessions from a wake.
  7. Preserve same-process uncollected media across continuation; make collection generation-scoped. Add an indexed parent-scoped pending scan and document the session v63 mixed-binary boundary.

Actual CI failure investigation

The original test job failed at TestTelegramSessionMgrResetSessionIfCurrent_CancelsActiveStream: cleanup calls = 0, want 1. This was inspected from the failed job log, not dismissed as unrelated. Race stress reproduced the failure locally: the stream finishes before deliberately deferred runner-owned cleanup. The test now synchronizes with the cleanup callback rather than goroutine scheduling. The same stress command passes after the fix (2,000 repetitions at each of 1/2/8/32 CPUs = 8,000 cases).

Validation on the follow-up

Go 1.26.8, temporary HOME/XDG and temporary test databases; module cache preserved:

  • go test ./... — passed.
  • go test -race ./cmd ./internal/tools ./internal/session ./internal/llm -run 'Test(Agent|PendingAgent|NativeAgent|HostLifecycle|ClaudeBinDelivers|MarkedLifecycle)' -count=1 — passed.
  • Additional explicit-stop/stopped-host-wake race regressions — passed.
  • go test -race ./internal/serve -run '^TestTelegramSessionMgrResetSessionIfCurrent_CancelsActiveStream$' -count=2000 -cpu=1,2,8,32 — passed after reproducing failures before the fix.
  • go vet ./..., make build, make complexity, git diff --check — passed.

Nested terminal modules and frontend source are unchanged; no separate local nested/frontend suites or live model/browser/Telegram delivery tests were run. Provider checks are serializer tests, not credentialed live requests.

Current CI

Head ec534ba094821625349dea4dd870be3fe01fa2a2, workflow https://github.com/SamSaffron/term-llm/actions/runs/37422156006, observed 2026-10-06 06:18 UTC:

  • test, changes, passkey-smoke, race, cross-build: passed.
  • frontend: failed at production-shaped Hub browser smoke: five node-order cases in frontend/e2e/hub-node-order.spec.ts (desktop drag/sidebar/node-token order and mobile/iPhone long-press order) time out with expected nodes absent. The failed job logs were fetched and inspected. Causality has not been established and this is not dismissed as unrelated. Reproduction/fix remains outstanding; recovery coordinator execution budget is exhausted.

Remaining caveats

  • Persisted-web-only host reactivation; unsupported surfaces reject rather than silently promise Telegram/CLI support.
  • Generation checks and reload ownership prevent the reproduced duplicate-delivery cases, but a hard crash between model/tool effects and acknowledgment cannot guarantee exactly-once effects. Recovery does not automatically restart children.
  • No automatic standalone browser push subscription is created by a host wake; replies use the existing web response/event channel.
  • Uncollected in-process child entries remain retained until collection; shutdown/runtime-store operations retain their existing cancellation limitations. These lower-priority review observations are not represented as fixed.

Address the existing Astra and Opus reviews of PR SamSaffron#1211. Preserve normal parent tools/settings, synthetic internal provenance, durable admission and reload acknowledgments, explicit stop intent, and uncollected media. Keep native children free of execution deadlines; 0–3600 second waits bound caller collection only and the configured/default 500 turns remain the execution budget. Synchronize the reproduced Telegram deferred-cleanup test race.
@SamSaffron
SamSaffron merged commit a7be981 into SamSaffron:main Oct 6, 2026
5 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants