Hi, AI review found a bug in parallel autovacuum (1ff3180ca01). I checked it myself and it reproduces on HEAD and PG19. Patch and reproducer attached.
When no DSM segment can be created, InitializeParallelDSM() does not fail. It sets up the parallel context in the leader's own memory with zero workers and leaves pcxt->seg NULL. Parallel query checks that pointer before using it, see ExecInitParallelPlan(), and VACUUM (PARALLEL) never looks at it and vacuums all the indexes in the leader. Parallel autovacuum registers a DSM detach callback on it, so the autovacuum worker segfaults and the database goes through crash recovery. It also leaves pv_shared_cost_params pointing into the leader's private memory with nothing left to reset it. It is not easy to hit. Right after the fallback, vacuum creates the DSA area for its dead items, and with no segments left that fails with a plain "too many dynamic shared memory segments" error before the callback is reached. Another backend has to free a segment in the window between the two. The fix is to set up the cost parameters and the callback only when the context has workers. The attached script fills the segments with parallel queries and keeps autovacuum busy on tables needing parallel index vacuuming. Most attempts fail with that error, then one lands in the window and crashes the worker, usually within a few seconds. [1] 2026-09-26 21:31:55.790 UTC [7908] LOG: autovacuum worker (PID 8904) was terminated by signal 11: Segmentation fault 2026-09-26 21:31:55.790 UTC [7908] DETAIL: Failed process was running: autovacuum: VACUUM ANALYZE public.t3 #0 slist_push_head (head=0x38, node=0x10468ff8) at ../../../../src/include/lib/ilist.h:1008 #1 on_dsm_detach (seg=0x0, function=0x775bce <parallel_vacuum_dsm_detach>, arg=0) at dsm.c:1148 #2 parallel_vacuum_init (rel=0x7f39f5bf0cf0, indrels=0x104be9a0, nindexes=3, nrequested_workers=2, vac_work_mem=65536, elevel=13, bstrategy=0x104ec410) at vacuumparallel.c:470 #3 dead_items_alloc (vacrel=0x104be050, nworkers=2) at vacuumlazy.c:3471 #4 heap_vacuum_rel (rel=0x7f39f5bf0cf0, params=0x7fffcadde380, bstrategy=0x104ec410) at vacuumlazy.c:863 #5 table_relation_vacuum (rel=0x7f39f5bf0cf0, params=0x7fffcadde380, bstrategy=0x104ec410) at ../../../src/include/access/tableam.h:1802 #6 vacuum_rel (relid=16399, relation=0x10501360, params=..., bstrategy=0x104ec410, isTopLevel=true) at vacuum.c:2353 #7 vacuum (relations=0x10501480, params=0x104fd5e0, bstrategy=0x104ec410, vac_context=0x10501260, isTopLevel=true) at vacuum.c:635 #8 autovacuum_do_vac_analyze (tab=0x104fd5d8, bstrategy=0x104ec410) at autovacuum.c:3351 #9 do_autovacuum () at autovacuum.c:2535 #10 AutoVacWorkerMain (startup_data=0x0, startup_data_len=0) at autovacuum.c:1634 #11 postmaster_child_launch (child_type=B_AUTOVAC_WORKER, child_slot=10, startup_data=0x0, startup_data_len=0, client_sock=0x0) at launch_backend.c:268 #12 StartChildProcess (type=B_AUTOVAC_WORKER) at postmaster.c:4043 #13 StartAutovacuumWorker () at postmaster.c:4107 #14 process_pm_pmsignal () at postmaster.c:3864 #15 ServerLoop () at postmaster.c:1720 #16 PostmasterMain (argc=3, argv=0x1043c510) at postmaster.c:1414 #17 main (argc=3, argv=0x1043c510) at main.c:227 -- Bharath Rupireddy Amazon Web Services: https://aws.amazon.com
From 9fb388bc6eed50ed56d39f88e5a73764db48c812 Mon Sep 17 00:00:00 2001 From: Bharath Rupireddy <[email protected]> Date: Sat, 26 Sep 2026 21:39:40 +0000 Subject: [PATCH v1] Fix crash in parallel autovacuum when no DSM segment can be created. When the maximum number of DSM segments has been reached, setting up a parallel context does not fail. It falls back to the leader's private memory and reports zero workers, and callers are expected to cope with that. Parallel query and manual VACUUM do, but the autovacuum part of parallel vacuum registered a segment detach callback without checking for the fallback. That crashes the autovacuum worker and takes the database through crash recovery. It also left the shared cost parameters pointing into the leader's private memory with nothing left to reset that pointer, so an error while vacuuming the table would leave a stale one behind for the next table. Hitting this is a bit difficult. The dead item store created right after the fallback normally fails with "too many dynamic shared memory segments" once the DSM segment limit is reached, so another backend has to free a segment in the small window between the two. Fix this by setting up the shared cost parameters and the detach callback only when the parallel context has workers, as there are no workers to propagate the parameters to otherwise. Oversight in commit 1ff3180ca01. Reported-by: Claude Code Author: Bharath Rupireddy <[email protected]> Discussion: https://postgr.es/m/<<message-id>> Backpatch-through: 19 --- src/backend/commands/vacuumparallel.c | 8 ++++++-- 1 file changed, 6 insertions(+), 2 deletions(-) diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c index 4532da60c84..d4572861000 100644 --- a/src/backend/commands/vacuumparallel.c +++ b/src/backend/commands/vacuumparallel.c @@ -458,9 +458,13 @@ parallel_vacuum_init(Relation rel, Relation *indrels, int nindexes, /* * Initialize shared cost-based vacuum delay parameters if it's for - * autovacuum. + * autovacuum and the parallel context has workers. Note that the parallel + * context falls back to the leader's private memory with no workers and + * no segment when the maximum number of DSM segments has been reached. + * There are then no workers to propagate the parameters to, and no + * segment to register the detach callback on. */ - if (shared->is_autovacuum) + if (shared->is_autovacuum && pcxt->nworkers > 0) { parallel_vacuum_set_cost_parameters(&shared->cost_params); pg_atomic_init_u32(&shared->cost_params.generation, 1); -- 2.47.3
repro.sh
Description: Bourne shell script
