Skip to content

Phase 8b: async-runtime adapter — the executor pool and its timer - #93

Merged
Wahbeh-Mohammad merged 5 commits into
mainfrom
31-phase-8b-async-runtime-adapter
Sep 22, 2026
Merged

Wahbeh-Mohammad merged 5 commits into
mainfrom
31-phase-8b-async-runtime-adapter

Conversation

@Wahbeh-Mohammad

Copy link
Copy Markdown
Contributor

Part of #31. First PR of phase 8b's three-PR stack — the code — built off main at a7cfeb6 (phases 0–7 and 8a) concurrently with 8c #32, reviewed and approved there; main has not moved since, so no reconcile pass was needed. Its tests are the next PR up and the phase record the one after. The first executor Transport.async_over has ever had, and the first real producer to settle phase 2's pivot.

What lands

dexpace-async-thread becomes a real gem — 10 files, +1,053 / −14 — ASYNC-1, 2, 5, 7–20 (seventeen ✅), ASYNC-3 ⏳ and ASYNC-4 N/A under §10.5 exactly as the issue forecast, plus fifteen cross-reference rows (PIPE-33, ASYNC-6, SEAM-12, SEAM-18, SEAM-25, XCUT-11, XCUT-13, XCUT-22, …).

  • Dexpace::Async::Thread::Pool — a fixed-size ::Thread pool over a bounded ::Thread::SizedQueue (size: required, queue_limit: derived from QUEUE_DEPTH_PER_WORKER), .build validating through four private class methods, #initialize keyword-only. #post never blocks the caller: a full queue or a closed pool raises RejectedError synchronously, which Bridge::AsyncOver routes to Completer#fail — the rejection policy that satisfies ASYNC-2 rather than violating it (R12). The worker loop is named, netted (rescue ::Exception → one contained http.instrumentation.hook ERROR diagnostic; report_on_exception = false), and performs the ASYNC-8–ASYNC-12 diagnostic hop: the caller's Fiber[] snapshot captured at #post, installed for the task through Diagnostics.with, and cleared at both boundaries — after every task and before the next — so nothing a task writes leaks to the next job on the same worker (P8-20). The gem's own defect diagnostic is emitted inside the installed context so it carries the caller's trace.id — round 4's R3-1, the one real code defect the reviews found: the report had been emitted after Diagnostics.with restored the thread, on the worker and on the timer alike.
  • #close / #release — Closeable's latch; the queue closed, the sentinels drained within shutdown_timeout: (DEFAULT_SHUTDOWN_TIMEOUT), the workers joined inside what remains so Thread.list is stable at return, SEAM-25's INSTRUMENTATION_SHUTDOWN emitted once at INFO with worker_count and drained (ASYNC-15–17). A close from one of the pool's own threads completes instead of self-joining (round 1's R0-2/R0-8: drain_workers counts the calling worker as exited; Timer#stop never joins ::Thread.current).
  • #delay(seconds) over the private Timer (ASYNC-18, R11): a parked timer thread over the pool's clock: and a ::Thread::Queue#pop(timeout:) wake — never Kernel#sleep, never a Fiber.scheduler; the future settles with true (P8-71 — Completer#fulfil(nil) raises; 5a's P5-52 precedent), a zero delay settles inline, a NaN, infinite or Complex duration is refused at the call (round 2's R1-1 — a NaN had killed the timer thread). Every callback runs guarded under the delay caller's own snapshot (round 3's R2-1 — the timer had inherited the first caller's Fiber[] for every later delay), entries are removed by identity, #schedule after #stop is refused through on_shutdown, and the timer carries the worker's net with a bounded join that never re-raises a dead thread's exception past #close (P8-22 and P8-25 extended, P8-75). private_constant :Timer, with a hooks.rbs-style sig mirror.
  • The entry file — require "dexpace", REQUIRED_CORE (7a's spelling, P8-72; the design's CORE_REQUIREMENT renamed, never aliased) and the version-skew assertion run before the require_relative chain, so a skewed Dexpace::VERSION fails the require itself and defines no Pool — the direct form P8-21 chose because there is no executor registry (P2-1). The gemspec is byte-identical: dexpace-core and nothing else.
  • sig/ mirrors lib/ one file per file; the manifest regenerated once — dexpace-async-thread 2 → 14 rows (Pool, .build, #post, #delay, #size, #queue_limit, #name, the three constants, REQUIRED_CORE, RejectedError; nothing for the private Timer, Job or Entry); the other five manifests byte-identical to main. The skeleton's smoke suite preloads core's stdlib before its constant snapshot and pins the four public constants.

Touches nothing else: no core file (Bridge::AsyncOver already checks the token before dispatch — phase 2's #44 — so the design's finding 2 was closed before this phase, and two of the plan's tests were rewritten to assert what core does), no tool, task, gate, Steepfile, rbs_collection.yaml, VERSIONS, Gemfile or CI file, nothing under 8c's, 8a's or 7a's gems. The plan's Task 1 allowlist fix and Task 12 composed clean-bundle check were both already on main (phase 0's #36; 8a's clean_bundle_check installs core beside every adapter) and are recorded as verification passes.

Decisions taken in the open, against the plan's text

Ledger rows P8-71–P8-78 in the design's As-built addendum and the checklist's "Deviations from the plan" (forty-one items). Beyond those above: the composed test at the charter's convergence point 2 is this lane's (8a landed first) and lives in this gem's test tree (P8-73 — see the tests PR); worker-side reads of the diagnostic context go through Diagnostics.capture because the 3.2 floor retains nil-valued Fiber[] keys, and core's reserved dexpace. slots travel with the snapshot (P8-74); the plan's Task 7 Step 8 "deadlock proof" is measured false (the timer's wait runs on the timer thread, which has no scheduler) and routed; the doubles are unique top-level Pool-prefixed constants and the test classes are gem-unique too (PoolMatrixFactsTest — the one-process test:gems run aborted at load on a bare MatrixFactsTest), now a CLAUDE.md constraint.

Layering

Each tip is green under every gate on its own tree — all eighteen. This branch is green on the SimpleCov floor too — 98.43% on 4.0.6 (3,704 runs, 0 failures); the tests PR takes the same tree to 99.87% with 3,818 runs / 73,279 assertions and 1 skip — 8a's TRANSPORT-18 measured vacuity, on main today.

Verification

  • Independent review, five rounds by five fresh reviewers with a fix round between each — the four-review cap was reached with round 3 at 0 blocking / 1 should-fix / 1 nit, and the manager chose one targeted round because R3-1 was a real code defect (6c's and 8a's precedent): round 0 1 / 2 / 5; round 1 0 / 1 / 2; round 2 1 / 0 / 2; round 3 0 / 1 / 1; round 4 approve, 0 / 0 / 0. Code changes from review: a close from the pool's own threads (round 1), the non-finite duration (round 2), the timer's per-delay context (round 3), the correlated defect diagnostic (round 4); every other finding was a test or docs gap, each closed by a mutation-proven case.
  • All eighteen gates individually at this tip on 4.0.6; the matrix set on 3.2.11 (98.48%), and on 3.3.12 and 3.4.10 at the tests tip; honest RuboCop clean (672 files); probe clean at the docs tip — each run by the reviewers and again by the manager's pre-push checks.
  • Mutations: 38, 43, 52, 47 and 42 by the five reviewers on 4.0.6 and 3.2.11; the one survivor at the last round is the recorded equivalent (guard 8 alone — the closed? pre-check in #post, whose red proof is 8 + 5 together). Forty-three guards run red and recorded in the checklist.
  • By experiment, both interpreters: a block posted under {trace.id: REQ-7} that raises produces a diagnostic carrying REQ-7 (nil at the round-3 tip); a close from a #delay future's on_settle returns with the shutdown event emitted; a stuck on_settle handler cannot hang #close; #delay(Float::NAN) refused synchronously; the real Dexpace::Transport::NetHTTP over WireServer through Transport.async_over(adapter, executor: pool) — a 200 as the exact object, a scripted failure as the same exception, an in-window cancel closing the pump-backed response exactly once; the as-built page's eleven blocks extracted and run as one script on both rows (56 checks).

Known follow-ups from the reviews (not blocking a gate)

  • Recorded equivalents: guard 8 alone (the closed? pre-check in #post) and guard 12 alone (return false if budget <= 0) — each survives alone; 8 + 5 and 12b are the red proofs.
  • #on_settle on a #delay future runs on the timer thread, or on the closing thread when #close fails it; a stuck handler there is the caller's own doing — documented in the README's "Where callbacks run", not guarded against. The shutdown callback runs under the delay caller's context on the closing thread and is restored, never cleared (P8-78).
  • Routed to phase 10's inbound list, dated: under the one-process test:gems the main fiber's storage held 5c's NO_SPAN slot when this gem's diagnostics suite ran (whether DexpaceTestCase should assert the main fiber's storage unchanged at teardown, and which suite leaves it); the plan's Task 7 Step 8 premise (whether XCUT-11's audit wants a structural scan for a mutex held across a queue wait).
  • pool_delay_test.rb and the as-built page reach the timer's entry list through instance_variable_get; a rename of those ivars needs the tests updated. Guard 5's racing test catches the missing ClosedQueueError arm only when the window opens; the deterministic "queue closed under an open latch" case is the load-bearing guard.

Build the async-runtime adapter: Dexpace::Async::Thread::Pool, a fixed-size
::Thread pool over a bounded ::Thread::SizedQueue that is the first real
implementation of SEAM-18's caller-supplied executor duck type and of
Dexpace::Page::_Executor, settles phase 2's core-owned pivot from a worker
thread, carries ASYNC-15..ASYNC-17's lifecycle over Dexpace::Closeable with
SEAM-25's shutdown event, performs the ASYNC-8..ASYNC-12 diagnostic hop
over Fiber[], and ships ASYNC-18's scheduled delay on one timer thread.

Pool.build(size:, ...) creates exactly size workers and never grows; #post
never blocks the caller (a full queue is RejectedError, a closed pool
Dexpace::ClosedError, both translated from the queue's own errors at the
one call site and routed to the failure channel by phase 2's bridge,
P8-23); #delay answers a Dexpace::Async::Future settled with true, as 5a's
Async.delay does, because Completer#fulfil(nil) raises (P8-71); #close
closes the queue, stops the private Timer and fails its outstanding
delays, drains the worker exit sentinels and joins the workers within one
shutdown_timeout budget, then emits Events::INSTRUMENTATION_SHUTDOWN once
(P8-24). A worker clears its inherited fiber storage at thread start and
again after every task so Diagnostics.with installs the caller's snapshot
rather than merging it onto the pool builder's or the previous task's
(P8-20); both threads the pool owns rescue ::Exception and report the
defect as a contained diagnostic instead of dying (P8-22, extended to the
timer, whose #schedule after #stop refuses the entry, P8-75). The entry
file asserts Dexpace::VERSION against REQUIRED_CORE at require time,
because SEAM-18 leaves no executor registry to register into (P8-21,
P8-72). Timer is a private_constant with a hooks.rbs-style sig mirror.

The smoke suite preloads core before its snapshot and pins the four public
constants; the surface manifest gains the object model's twelve rows and
no private one. The gemspec is unchanged: dexpace-core and nothing else.
…ining

Phase 8b, review round 0's R0-2 and R0-8. A Pool#close issued from a
#delay future's on_settle handler -- which runs on the timer thread, as
the README documents -- reached Timer#stop's thread.join on the current
thread: ThreadError escaped #close with the latch already flipped, no
shutdown event was emitted and every other outstanding delay stayed
unsettled forever (measured on 4.0.6 and 3.2.11). Timer#stop now skips
the join when the thread is the current one; the wake queue is closed
regardless, so the next pop answers nil and #run exits as soon as the
handler returns, and the leftovers are failed on the closing thread as
on any other.

The same self-wait, one thread over: a #close issued from inside a
posted task waited for its own worker's exit sentinel, which cannot be
pushed until #release returns, so every such close burned the whole
shutdown budget and reported drained: false. drain_workers now counts
the calling worker as exited and never joins it; that worker exits by
itself once the task returns, draining whatever the closed queue still
holds first. Recorded as ledger row P8-76 on the docs branch; the YARD
on #delay and #release states both outcomes.
Phase 8b, review round 1's R1-1. Pool#delay validated a duration as
Numeric and not negative, and a NaN answers false to both negative? and
zero?: it reached the timer, whose list is ordered by deadline. The NaN
deadline made next_wait's max raise inside the timer thread's net -- one
diagnostic, the thread gone for good with @thread still set, the future
never settled -- and every later #delay on that pool raised a bare
ArgumentError synchronously from sort_by! (measured on 4.0.6 and
3.2.11). A Complex is a Numeric with no order and no #negative?, a bare
NoMethodError from the same method.

validate_duration! now asks real? before negative? and finite? after
it, so a NaN, an infinite or a Complex duration raises
Dexpace::InvalidArgumentError before anything is scheduled (ASYNC-18's
"MUST reject", P8-77). finite? is Numeric's own protocol and covers a
Float or BigDecimal NaN or infinity with one call, which is why Infinity
is refused with NaN: an entry that never fires is not a delay. The same
shape guarded shutdown_timeout: a NaN budget made #release's deadline
arithmetic raise a bare ArgumentError out of #close with the latch
already flipped, and an infinite one is the unbounded close XCUT-13
forbids for this gem by name, so non_negative_numeric! becomes
finite_non_negative! (real, non-negative and finite), renamed in the sig.
Phase 8b, review round 2's R2-1. The timer thread is the gem's second
::Thread.new carrier, and it inherited the FIRST #delay caller's fiber
storage at creation and was never cleared and never handed a per-delay
snapshot: every later delay's #on_settle and #then callback ran under a
stale, foreign context -- caller B's handler and a contextless caller
C's both read caller A's trace id, and a key one handler wrote was
visible to every later one (measured on 4.0.6 and 3.2.11). ASYNC-8
names callbacks explicitly, and the design's own finding 1 names this
mechanism as ASYNC-10's stale assembly-time snapshot; only the workers
had the floor P8-20 states.

Timer#schedule now captures Diagnostics.capture per entry, on the
scheduling caller's thread inside Pool#delay (ASYNC-10's per-submission
point); the Entry carries the snapshot; the timer thread clears its
inherited storage once at start and again after every fired callback,
which runs under Diagnostics.with over the entry's snapshot (ASYNC-8,
ASYNC-9) -- the shape Pool#run already has, applied to the second thread
the gem owns (P8-78, extending P8-20). The shutdown callback runs on
whichever thread stopped the timer under the same install-and-restore,
so the closer's own context is put back and never cleared: that thread
is not the timer's to empty. The sig declares the new member and the
three private methods; no public surface changes.
The gem's one log event emitted on a caller's behalf after the hop --
P8-22's ERROR http.instrumentation.hook for a posted block or a delay
handler that raised -- was emitted after Diagnostics.with had restored
the thread, so on the worker and on the timer alike it carried no
trace id (payload keys exactly [cause, event] on 4.0.6 and 3.2.11),
while a line the block itself logged inside the hop carried the id.
ASYNC-8's stated purpose is that events emitted after the hop retain
correlation, and this was the one event the gem itself emits there
(review round 3's R3-1).

Pool#run now rescues inside the Diagnostics.with block, so
report_failure runs with the task's snapshot still installed and the
ensure clear stays at method level; Timer#fire and Timer#shut_down
nest guarded inside Diagnostics.with rather than around it, so the
timer's on_error runs under the delay caller's snapshot on the timer
thread and on the closing thread. Nothing outside either net can
raise -- .with's install and restore are Fiber[]= over the Symbol keys
.capture read off a fiber's storage -- so both threads' survival stays
structural. The sig files are unchanged.
@Wahbeh-Mohammad Wahbeh-Mohammad added type:feature New capability or enhancement area:transport Transport, async model, seams: TRANSPORT-* ASYNC-* SEAM-* labels Sep 22, 2026
@Wahbeh-Mohammad
Wahbeh-Mohammad merged commit f2c4df4 into main Sep 22, 2026
5 checks passed
@Wahbeh-Mohammad Wahbeh-Mohammad mentioned this pull request Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:transport Transport, async model, seams: TRANSPORT-* ASYNC-* SEAM-* type:feature New capability or enhancement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant