Skip to content

Resume agent turns after auto-approval, and fix retryAgent and Scheduler docs - #599

Merged
ndisidore merged 4 commits into
mainfrom
fix/small-kernel-findings
Sep 30, 2026
Merged

ndisidore merged 4 commits into
mainfrom
fix/small-kernel-findings

Conversation

@ndisidore

@ndisidore ndisidore commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Resume agent turns that auto-approval decides

A turn paused on an awaited action resumed only when the user clicked Approve on that action. "Always approve" and the drain that runs after a manual approval applied actions without resuming the turn, so always-approving the first Gmail archive left the agent idle until the user sent another message. Both paths now resume the turn, and a drain requested while another is running waits for that one to finish, so the resume sees the applied action.

Fix the Scheduler's restore guidance for agents

The Scheduler's agent-facing types told agents to call ctx.restore() from executeCode. That call targets the executeCode worker and throws because the worker has no [restore]() method. The types and README now use env.MY_GADGET[restore](params), and the example calls env.SCHEDULER instead of the Workers global scheduler.

Document that retryAgent throws while an agent is running

The retryAgent doc said the call does nothing when an agent is already running. It throws the agent-running error, as the other chat methods do, and the frontend already shows a toast for it.


Devin Review

The implementation rejects with the agent-running error, as the other chat mutators do; the doc
said it silently did nothing.
@github-actions github-actions Bot added kernel Changes to the Workshop kernel gatekeeper Changes to a gatekeeper integration workshop/shared Changes to shared Workshop APIs labels Sep 29, 2026
@github-actions

Copy link
Copy Markdown

Preview: pr599-fix-small-ker-bab1b92c

https://pr599-fix-small-ker-bab1b92c-router.cloudflare-os-previews.workers.dev

Dashboard · deleted when this PR closes

ask-bonk[bot]

This comment was marked as resolved.

@ask-bonk

ask-bonk Bot commented Sep 29, 2026

Copy link
Copy Markdown

Review: 1 findings.

I posted one inline finding: auto-approval can finish before the submitting agent turn has suspended, leaving the turn idle instead of resuming it.

github run

devin-ai-integration[bot]

This comment was marked as resolved.

@ask-bonk

ask-bonk Bot commented Sep 29, 2026

Copy link
Copy Markdown

Review: 0 findings.

No new actionable issues beyond the existing inline discussions about auto-approval resumption.

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Sep 29, 2026
@ndisidore
ndisidore force-pushed the fix/small-kernel-findings branch from c36bd80 to 46d8e44 Compare September 29, 2026 21:43
ask-bonk[bot]

This comment was marked as resolved.

@ask-bonk

ask-bonk Bot commented Sep 29, 2026

Copy link
Copy Markdown

Review: 1 findings.

I posted one actionable inline finding.

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Sep 29, 2026
@ndisidore
ndisidore force-pushed the fix/small-kernel-findings branch from 46d8e44 to d6cd3d8 Compare September 29, 2026 22:29
@ask-bonk

ask-bonk Bot commented Sep 29, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Sep 29, 2026
@Maximo-Guk

This comment was marked as outdated.

The Scheduler's agent-facing types told agents to call ctx.restore() from executeCode, which
targets the executeCode worker and throws that it implements no [restore](). The example also
called a bare scheduler, the Workers global, rather than env.SCHEDULER.
@ndisidore
ndisidore force-pushed the fix/small-kernel-findings branch from d6cd3d8 to 735e772 Compare September 30, 2026 18:49
@ask-bonk

ask-bonk Bot commented Sep 30, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions

Copy link
Copy Markdown

Eval results

Verdict: ⚪ Unchanged. No task moved beyond what 10 runs can tell apart from noise.

Task Score Δ score Fisher test Cache hits Avg min Avg steps
change-calendar 78% (1 run error) → 100% baseline run errors — 83% 3.0 → 2.8 21.3 → 24.0
chess 90% 0 pp p = 1.00 96%
0 pp
8.1 → 7.3 56.9 → 55.7
incident-desk 100% → 90% −10 pp p = 1.00 93%
0 pp
3.9 → 3.4 32.0 → 29.4
worker-logs 100% → 80% −20 pp p = 0.47 91% → 90%
−1 pp
3.4 → 3.0 25.7 → 23.6
Failed checks
Task Check Failed
change-calendar t4 names-the-booked-windows-the-new-rules-reject 1/9 → 0/10
change-calendar t4 earliest-available-applies-every-rule 1/9 → 0/10
change-calendar t4 earliest-available-matches-the-reference 1/9 → 0/10
change-calendar t4 earliest-available-slots-are-bookable 1/9 → 0/10
chess t1 agrees-with-the-oracle-on-the-hard-positions 0/10 → 1/10
chess t1 agrees-with-the-oracle-on-perft-positions 0/10 → 1/10
chess t1 agrees-with-the-oracle-through-random-games 0/10 → 1/10
chess t2 imports-known-games-to-the-right-positions 1/10 → 0/9
chess t2 exports-pgn-the-oracle-replays-to-the-same-position 1/10 → 0/9
incident-desk t1 opens-acknowledges-and-resolves-in-order 0/10 → 1/10
incident-desk t1 simultaneous-acknowledges-yield-exactly-one-owner 0/10 → 1/10
incident-desk t1 simultaneous-opens-of-one-id-admit-exactly-one 0/10 → 1/10
incident-desk t1 board-lists-by-severity-then-age 0/10 → 1/10
worker-logs t2 filters-combine-across-colo-route-and-worker 0/10 → 2/10
worker-logs t2 existing-data-and-summary-survive 0/10 → 1/10
worker-logs t2 reset-and-ingest-still-work 0/10 → 1/10

Run · trajectories and raw results

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Sep 30, 2026
@ask-bonk

ask-bonk Bot commented Sep 30, 2026

Copy link
Copy Markdown

🔬 Eval runs review

Performance

Comparable tasks scored chess 90% → 90%, incident-desk 100% → 90%, and worker-logs 100% → 80%; comparison.json classifies these changes as unchanged, within ten-run noise. Change-calendar reached 10/10 versus main’s 7/9 completed runs, but its baseline stream error prevents a clean comparison. Worker-logs lost the most candidate runs to mistakes introduced while extending generated code; chess consumed the most steps and editing retries, with unmatched edits rising from 15 to 30 despite average steps falling from 56.9 to 55.7.

⚪ VERDICT: NO REGRESSION FROM THIS PR

The observed failures are model mistakes in generated gadgets, not consequences of the diff’s auto-approval resumption, Scheduler guidance, or retryAgent documentation changes, and no comparable pass rate moved beyond noise.

Triage

Failure modes

  • Query invariants lost during extension · worker-logs 2/10 · model error · this PR: no — In turn 2, trial 1’s writeFile(server.js) removed the query’s duration ordering without sorting before selecting p95, producing incorrect latencies despite retaining events; trial 10’s same call computed the time range but omitted it from makeFilters, so filtered summaries included requests outside the requested interval. Main’s passing implementations preserved ordering or sorted durations explicitly and retained range filtering.
  • Sliding-piece attacks detected only at adjacent squares · chess 1/10 · model error · this PR: no — Trial 2, turn 1’s writeFile(server.js) required d === 1 for rook/bishop attacks, admitting pinned moves and castling through distant rook attacks. The final reply claimed full rules without execution checks; main passed these rules checks in every run, although one main run separately failed PGN import/export in turn 2.
  • Empty SQL lookup treated as a nullable result · incident-desk 1/10 · model error · this PR: no — Trial 2, turn 1’s writeFile(server.js) used .one() for the duplicate-ID lookup before inserting a new incident; the recorded “got no results” exception prevented opening any incidents. Main’s passing code used .toArray()[0] for optional lookups; the candidate ended after five model steps without exercising open.

Tool errors

  • editFile: No matching text was found in the file. · change-calendar 2 → 0, chess 15 → 30, incident-desk 2 → 0, worker-logs 1 → 1 · model error — Agents supplied text differing from the file, notably overescaped regexes and newline literals during chess PGN edits; candidate trial 10 repeatedly reread and retried, accumulating 15 tool errors while eventually passing.
  • editFile: Validation failed · change-calendar 1 → 3, chess 1 → 5, incident-desk 4 → 1, worker-logs 3 → 2 · model error — Malformed arguments omitted the required replacement property; candidate change-calendar trial 1 sent .replacement, then corrected it on the next step.
  • editFile: Multiple matches were found. The text to match must be unique. · chess 1 → 3, incident-desk 7 → 3, worker-logs 1 → 2 · model error — Agents selected repeated snippets despite the parameter’s exactly-one-match requirement; candidate chess trial 10 tried replacing every this.serialized( and this.readPosition() occurrence, then searched for distinguishing context.
  • readFile: File does not exist. · change-calendar 2 → 3, incident-desk 2 → 0, worker-logs 4 → 2 · model error — Agents attempted to read server.js and client.js immediately after creating an empty gadget, despite createGadget documenting that default; candidate change-calendar trial 4 recovered by writing the files.

Prompt cache

No comparable task’s cache-hit rate moved by five percentage points or more; no task meets the requested investigation threshold.

What to do

  • Add focused tasks under packages/workshop-evals/evals/ for enabling auto-approval on a suspended turn and registering a Scheduler callback through env.MY_GADGET[restore]; these runs do not demonstrate the behavioral benefits targeted by this PR.
  • Optional, unrelated to this PR: add guidance in packages/workshop-backend/src/agent.ts to exercise RPC boundary cases through executeCode before claiming completion, including empty SQL lookups, distant chess attacks, and preservation of query semantics after extensions.

github run

Only a manual approveAction resumed a turn suspended on awaitDecision. "Always approve" (the
chat card, Activity panel and Connections rule toggle) and approveAction's cascade apply actions
through the drain alone, so when the drain decided a turn's last awaited action -- e.g. the first
Gmail archive or a vetted MCP tool call -- the action applied and the agent never resumed or
learned of it.

Both drains now resume the chats whose awaited actions were pending on that gatekeeper. A drain
requested while one is running now shares its promise, so awaiting drain() waits for the rerun
it requested instead of returning before the action is applied. No caller awaits a drain, so a
failed resume is logged rather than ending the loop.

An auto-approvable awaited action queued behind a manual gate, which the in-order drain stops at,
also let its turn run on while the action sat unapplied. submitAction now suspends that turn too,
and rejectAction drains as approveAction does, so rejecting the gate applies the queued action
and resumes the turn instead of leaving both waiting on a drain that nothing starts.
The drain-and-resume path checked every chat with an awaited action pending on the gatekeeper,
including ones the drain left pending, so a chat whose older action stayed on a manual gate could
be resumed from its latest turn's already-approved actions.
@ndisidore
ndisidore force-pushed the fix/small-kernel-findings branch from 735e772 to 35de05d Compare September 30, 2026 19:33
@ask-bonk

ask-bonk Bot commented Sep 30, 2026

Copy link
Copy Markdown

LGTM!

github run

@ndisidore
ndisidore merged commit 1200259 into main Sep 30, 2026
16 of 17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

gatekeeper Changes to a gatekeeper integration kernel Changes to the Workshop kernel workshop/shared Changes to shared Workshop APIs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants