Skip to content

Add residency-budget and worker-crash TAP tests - #59

Open
matt-welch wants to merge 2 commits into
mainfrom
v04-score-verification
Open

matt-welch wants to merge 2 commits into
mainfrom
v04-score-verification

Conversation

@matt-welch

Copy link
Copy Markdown
Contributor

Description

Adds two independent TAP tests that exercise runtime behavior which had previously only been reasoned about by reading the code:

  • test/t/50_undo_growth_bound.pl confirms that a single large transaction's undo-log growth stays a small fraction of its database's residency budget, both at the default budget and at a raised one, and that the stop point scales with the budget. It also confirms that capacity-headroom inserts (charged 0 bytes against the budget) are finite rather than an unbounded loophole in that gate.
  • test/t/51_worker_crash_scope.pl confirms that a per-database vamana worker's crash scope depends on how it died: a graceful exit (pg_terminate_backend/SIGTERM) stays contained to its own database, while a signal-killed shared-memory-attached process (the worker, a parked search-slot process, or the launcher) makes the postmaster treat it as a crash and restart every backend in the cluster.

No source files are changed by this PR. Both tests were run repeatedly against current main and passed every time with consistent results across runs; see "Testing Notes" below for exact tallies.

Related Issues

None.

Type of Change

  • Bug fix
  • New feature
  • Refactor / code cleanup
  • Documentation update
  • Test addition or update
  • Build / CI change

Pre-Merge Checklist

Build

  • make completes without errors or warnings
  • make install completes successfully

Tests

  • If this PR introduces no new behavior: existing regression tests (make installcheck) and TAP tests (test/t/) pass with no failures
  • If this PR introduces new behavior: test cases covering it were added to test/sql/ and/or test/t/
  • If this PR adds a standalone unit-test module under test/modules/: it builds and passes (make -C test/modules/<module> installcheck)

Documentation

  • Relevant docs under docs/ updated if architecture or usage changed

Testing Notes

Both new TAP files were built and run individually against current main, not as part of a full-suite run, so the "existing regression tests and TAP tests pass with no failures" checklist item above is left unchecked: that claim was not verified in this PR and should be confirmed by CI or a full-suite run before merge. No source files were touched, so the risk of regression is low, but it has not been directly checked here.

  • test/t/50_undo_growth_bound.pl: 20/20 tests passed. Three runs at the default 100MB residency budget each stopped at exactly 130,000 rows; three runs at a raised 200MB budget each stopped at exactly 260,000 rows (confirming the stop point scales with the budget); a larger-seed-index run confirmed capacity-headroom inserts are finite.
  • test/t/51_worker_crash_scope.pl: 24/24 tests passed. Three repeats of the graceful-termination case stayed contained to the single database each time; SIGSEGV, SIGABRT, and SIGKILL on the worker, plus SIGKILL on a parked search-slot process and on the launcher, each produced the cluster-wide crash-recovery sequence.

A single large transaction's undo-log growth is bounded two ways:
every INSERT reserves residency bytes against its database's budget
before the undo entry is appended, so a huge single-transaction
INSERT hits the residency-budget ERROR long before undo memory grows
large; and the shared pending array's repalloc growth is capped by
PostgreSQL's MaxAllocSize, stopping near 33.5M entries with a clean
ERROR rather than an OOM kill. This had been reasoned about by
reading the code only; this test exercises it directly.

- Inserts in batches inside one transaction until the residency
  budget ERRORs, at both the default and a raised per-database
  budget, confirming the stop point scales with the budget and undo
  memory stays under 3% of it either way
- Builds a larger seed index to check that capacity-headroom inserts
  (charged 0 bytes against the budget) are finite, not an unbounded
  loophole in the gate

Signed-off-by: Matt Welch <matt.welch@intel.com>
A per-database vamana worker's crash scope depends on how it died: a
graceful exit (pg_terminate_backend/SIGTERM) stays contained to its
own database, but a signal-killed shared-memory-attached process
(SIGSEGV, SIGABRT, SIGKILL) makes the postmaster treat it as a crash
and restart every backend in the cluster. This had been reasoned
about by reading the code only; this test exercises it directly.

 - Confirms the contained case three times: the other database's
   session survives with the same backend pid, no cluster-wide
   restart message appears, and the launcher respawns the killed
   worker with a new pid
 - Confirms every signal case (SEGV, ABRT, KILL) produces the
   cluster-wide crash-recovery sequence and kills the other
   database's session too
 - Extends the same signal test to a parked search-slot process and
   to the launcher process itself, confirming the cluster-wide crash
   path is not specific to the worker

Signed-off-by: Matt Welch <matt.welch@intel.com>
@matt-welch
matt-welch requested a review from a team October 1, 2026 20:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant