Hi,

After getting involved in the discussion of corruption caused by
CREATE DATABASE (from template) STRATEGY WAL_LOG [1], I started
thinking about how we could prevent and fix VM corruption in more
cases.

After ed62d26caca, the idea was that both setting and clearing the VM
always registered the VM block, so we would be protected against VM
corruption. This turned out to not work if the VM got removed somehow
from the primary or standby and they got out-of-sync (like in [1]). I
have some stop-gap fixes proposed for that in [2]. However, I think if
we make some bigger changes, we can prevent and repair scenarios like
this.

The patches attached do the following visibility map hardening:
- WAL-log page corruption repair so that it is propagated to the
standby (and isn't lost after a crash). Also extend page corruption
repair to cover a few more scenarios.
- Register the VM buffer whenever INSERT/UPDATE/DELETE clears
PD_ALL_VISIBLE, even if the VM bits were already clear on the primary.
This repairs divergence and can't result in torn pages during recovery
like the stop-gap strategy.
- Stop reading the VM with RBM_ZERO_ON_ERROR. Introduce a new mode,
RBM_ZERO_ON_MISSING, that allows extending if the page is not created
yet but errors on corrupt pages.

This series also makes zero_damaged_pages a string instead of a
boolean so you can specify 'vm' instead of just on and allow you to
ignore and zero out the VM if it is corrupt during recovery.

There's also a patch to stop masking PD_ALL_VISIBLE WAL consistency
checking during recovery in [3] (see v6-0002), which I plan to commit
soon once sources of this kind of corruption are fixed.

The aim is to prevent avoidable divergence, make repairs durable, and
leave better evidence when corruption does occur. (note that patch set
doesn't have tests yet)

- Melanie

[1] 
https://www.postgresql.org/message-id/CAAKRu_bK7oJvtrrzL_V-OHouCWDKyU2-ERv%2BQvV6XM9Wh2dpUg%40mail.gmail.com
[2] 
https://www.postgresql.org/message-id/CAAKRu_bApoksLDb-HX0GYciU3uLWqA1JagntaV8GP0%3D%2BidehHw%40mail.gmail.com
[3] 
https://www.postgresql.org/message-id/CAAKRu_aRdQ6RjneKeQh9%2BRVMFgMrtjTYFX%3DzDgBdSiQ8JDs6Jg%40mail.gmail.com
From 687fb98240d315082dc47687ad7ac6f119956909 Mon Sep 17 00:00:00 2001
From: Melanie Plageman <[email protected]>
Date: Thu, 24 Sep 2026 16:42:36 -0400
Subject: [PATCH v1 01/11] Detect and repair a stale all-frozen visibility map
 bit

visibilitymap_set() only ORs in the requested bits. If vacuum finds that
every tuple on a page is visible but not every tuple is frozen, setting
all-visible cannot repair an incorrectly set all-frozen bit. The function
marks the VM buffer dirty when the requested flags differ from the existing
bits, even if the OR does not change those bits, so dirtying the buffer
alone does not correct the corruption.

Add a VM_CORRUPT_STALE_ALL_FROZEN case to heap_page_fix_vm_corruption()
that warns and clears the VM bits when a full scan establishes that the
page will remain all-visible but not all-frozen. The normal path then sets
all-visible again without retaining the stale all-frozen bit.

Backpatch-through: 19
---
 src/backend/access/heap/pruneheap.c | 33 +++++++++++++++++++++++++++++
 1 file changed, 33 insertions(+)

diff --git a/src/backend/access/heap/pruneheap.c b/src/backend/access/heap/pruneheap.c
index 98fba4bb7c1..a8e17fcc8d2 100644
--- a/src/backend/access/heap/pruneheap.c
+++ b/src/backend/access/heap/pruneheap.c
@@ -196,6 +196,8 @@ typedef enum VMCorruptionType
 	VM_CORRUPT_LPDEAD,
 	/* Tuple not visible to all transactions on a page marked all-visible */
 	VM_CORRUPT_TUPLE_VISIBILITY,
+	/* Page marked all-frozen in the VM but not actually all-frozen */
+	VM_CORRUPT_STALE_ALL_FROZEN,
 } VMCorruptionType;
 
 /* Local functions */
@@ -944,6 +946,23 @@ heap_page_fix_vm_corruption(PruneState *prstate, OffsetNumber offnum,
 								relname, prstate->block)));
 			do_clear_vm = true;
 			break;
+
+		case VM_CORRUPT_STALE_ALL_FROZEN:
+
+			/*
+			 * We examined every tuple on the page and found that the page is
+			 * all-visible but not all-frozen, yet the VM marks it all-frozen.
+			 * The page-level PD_ALL_VISIBLE flag is still correct, so only
+			 * the VM is wrong. Clear both the VM bits. All-visible will be
+			 * set again through the normal path.
+			 */
+			ereport(WARNING,
+					(errcode(ERRCODE_DATA_CORRUPTED),
+					 errmsg("page marked all-frozen in the visibility map is not all-frozen"),
+					 errcontext("relation \"%s\", page %u",
+								relname, prstate->block)));
+			do_clear_vm = true;
+			break;
 	}
 
 	Assert(do_clear_heap || do_clear_vm);
@@ -1259,6 +1278,20 @@ heap_page_prune_and_freeze(PruneFreezeParams *params,
 	Assert(!prstate.set_all_visible || prstate.attempt_set_vm);
 	Assert(!prstate.set_all_visible || (prstate.lpdead_items == 0));
 
+	/*
+	 * If we examined every tuple on the page and found that the page is
+	 * all-visible but not all-frozen, yet the VM marks it all-frozen, that
+	 * all-frozen bit is corrupt. Repair the VM. We will set it back to
+	 * all-visible later along with the other changes. Note that this must be
+	 * done after set_all_visible and set_all_frozen are finalized above to
+	 * account for dead items and unfrozen tuples.
+	 */
+	if (prstate.attempt_freeze && prstate.set_all_visible &&
+		!prstate.set_all_frozen &&
+		(prstate.old_vmbits & VISIBILITYMAP_ALL_FROZEN))
+		heap_page_fix_vm_corruption(&prstate, InvalidOffsetNumber,
+									VM_CORRUPT_STALE_ALL_FROZEN);
+
 	do_set_vm = heap_page_will_set_vm(&prstate, params->reason, do_prune, do_freeze);
 
 	/*
-- 
2.43.0

From 5a9e4da830c9e628ef5d0ea30fa4f1bddfab3f3d Mon Sep 17 00:00:00 2001
From: Melanie Plageman <[email protected]>
Date: Fri, 25 Sep 2026 15:02:18 -0400
Subject: [PATCH v1 02/11] WAL-log visibility map corruption repair

When VACUUM finds the visibility map inconsistent with a heap page, it
repairs the divergence, but that repair was not WAL-logged. The corruption
therefore persisted on standbys and could resurface after a failover, and
the repair itself was not torn-page safe.

Log the repair as a single XLOG_FPI record carrying full-page images of
the heap page and/or VM page that were changed. VM corruption is rare,
so the extra WAL does not matter.
---
 src/backend/access/heap/pruneheap.c | 59 ++++++++++++++++++++++++++---
 1 file changed, 53 insertions(+), 6 deletions(-)

diff --git a/src/backend/access/heap/pruneheap.c b/src/backend/access/heap/pruneheap.c
index a8e17fcc8d2..ecfe9164032 100644
--- a/src/backend/access/heap/pruneheap.c
+++ b/src/backend/access/heap/pruneheap.c
@@ -22,6 +22,7 @@
 #include "access/visibilitymap.h"
 #include "access/xlog.h"
 #include "access/xloginsert.h"
+#include "catalog/pg_control.h"
 #include "commands/vacuum.h"
 #include "executor/instrument.h"
 #include "miscadmin.h"
@@ -878,8 +879,8 @@ heap_page_will_freeze(bool did_tuple_hint_fpi,
  * the heap buffer is exclusively locked, ensuring that no other backend can
  * update the VM bits corresponding to this heap page.
  *
- * This function makes changes to the VM and, potentially, the heap page, but
- * it does not need to be done in a critical section.
+ * This function makes changes to the VM and, potentially, the heap page, and
+ * WAL-logs them in its own critical section.
  */
 static void
 heap_page_fix_vm_corruption(PruneState *prstate, OffsetNumber offnum,
@@ -967,21 +968,67 @@ heap_page_fix_vm_corruption(PruneState *prstate, OffsetNumber offnum,
 
 	Assert(do_clear_heap || do_clear_vm);
 
-	/* Avoid marking the buffer dirty if PD_ALL_VISIBLE is already clear */
+	if (do_clear_vm)
+		LockBuffer(prstate->vmbuffer, BUFFER_LOCK_EXCLUSIVE);
+
+	START_CRIT_SECTION();
+
 	if (do_clear_heap)
 	{
 		Assert(PageIsAllVisible(prstate->page));
 		PageClearAllVisible(prstate->page);
-		MarkBufferDirtyHint(prstate->buffer, true);
+		MarkBufferDirty(prstate->buffer);
 	}
 
 	if (do_clear_vm)
 	{
-		LockBuffer(prstate->vmbuffer, BUFFER_LOCK_EXCLUSIVE);
-		/* This VM clear is not WAL-logged, so its return value is not needed. */
 		(void) visibilitymap_clear(prstate->relation->rd_locator,
 								   prstate->block, prstate->vmbuffer,
 								   VISIBILITYMAP_VALID_BITS);
+
+		/*
+		 * The VM bits might already be clear even though PD_ALL_VISIBLE was
+		 * incorrectly set. Still WAL-log the VM image as the caller reported
+		 * a type of corruption that would normally require clearing the VM.
+		 * Logging it ensures the VM is also cleared on the standby. If we log
+		 * it, we have to mark it dirty.
+		 */
+		MarkBufferDirty(prstate->vmbuffer);
+	}
+
+	/*
+	 * WAL-log the repair so that standbys and crash recovery apply the same
+	 * fix and the VM stays in sync across a cluster.
+	 *
+	 * Rather than inventing a dedicated record type, just log full-page
+	 * images of the pages we changed. VM corruption is rare, so the extra WAL
+	 * does not matter.
+	 */
+	if (RelationNeedsWAL(prstate->relation))
+	{
+		XLogRecPtr	recptr;
+		uint8		block_id = 0;
+
+		XLogBeginInsert();
+		if (do_clear_heap)
+			XLogRegisterBuffer(block_id++, prstate->buffer,
+							   REGBUF_FORCE_IMAGE | REGBUF_STANDARD);
+		if (do_clear_vm)
+			XLogRegisterBuffer(block_id++, prstate->vmbuffer,
+							   REGBUF_FORCE_IMAGE);
+
+		recptr = XLogInsert(RM_XLOG_ID, XLOG_FPI);
+
+		if (do_clear_heap)
+			PageSetLSN(prstate->page, recptr);
+		if (do_clear_vm)
+			PageSetLSN(BufferGetPage(prstate->vmbuffer), recptr);
+	}
+
+	END_CRIT_SECTION();
+
+	if (do_clear_vm)
+	{
 		LockBuffer(prstate->vmbuffer, BUFFER_LOCK_UNLOCK);
 		prstate->old_vmbits = 0;
 	}
-- 
2.43.0

From 664095ff205cb944693b149323d122ba97090eb4 Mon Sep 17 00:00:00 2001
From: Melanie Plageman <[email protected]>
Date: Fri, 25 Sep 2026 15:02:24 -0400
Subject: [PATCH v1 03/11] Make heap_page_fix_vm_corruption() usable outside of
 pruning

heap_page_fix_vm_corruption() took the PruneState, which confined it to
pruneheap.c. Have it take the relation, heap buffer and VM buffer instead,
export it, and move VMCorruptionType to heapam.h so that vacuumlazy.c can
use it too.
---
 src/backend/access/heap/pruneheap.c | 121 +++++++++++++++-------------
 src/include/access/heapam.h         |  21 +++++
 2 files changed, 84 insertions(+), 58 deletions(-)

diff --git a/src/backend/access/heap/pruneheap.c b/src/backend/access/heap/pruneheap.c
index ecfe9164032..bfd091781dd 100644
--- a/src/backend/access/heap/pruneheap.c
+++ b/src/backend/access/heap/pruneheap.c
@@ -183,33 +183,12 @@ typedef struct
 	OffsetNumber *deadoffsets;	/* points directly to presult->deadoffsets */
 } PruneState;
 
-/*
- * Type of visibility map corruption detected on a heap page and its
- * associated VM page. Passed to heap_page_fix_vm_corruption() so the caller
- * can specify what it found rather than having the function rederive the
- * corruption from page state.
- */
-typedef enum VMCorruptionType
-{
-	/* VM bits are set but the heap page-level PD_ALL_VISIBLE flag is not */
-	VM_CORRUPT_MISSING_PAGE_HINT,
-	/* LP_DEAD line pointers found on a page marked all-visible */
-	VM_CORRUPT_LPDEAD,
-	/* Tuple not visible to all transactions on a page marked all-visible */
-	VM_CORRUPT_TUPLE_VISIBILITY,
-	/* Page marked all-frozen in the VM but not actually all-frozen */
-	VM_CORRUPT_STALE_ALL_FROZEN,
-} VMCorruptionType;
-
 /* Local functions */
 static void prune_freeze_setup(PruneFreezeParams *params,
 							   TransactionId *new_relfrozen_xid,
 							   MultiXactId *new_relmin_mxid,
 							   PruneFreezeResult *presult,
 							   PruneState *prstate);
-static void heap_page_fix_vm_corruption(PruneState *prstate,
-										OffsetNumber offnum,
-										VMCorruptionType corruption_type);
 static void prune_freeze_fast_path(PruneState *prstate,
 								   PruneFreezeResult *presult);
 static void prune_freeze_plan(PruneState *prstate,
@@ -871,26 +850,34 @@ heap_page_will_freeze(bool did_tuple_hint_fpi,
  * The caller specifies the type of corruption it has already detected via
  * corruption_type, so that we can emit the appropriate warning. All cases
  * result in the VM bits being cleared; corruption types where PD_ALL_VISIBLE
- * is incorrectly set also clear PD_ALL_VISIBLE.
+ * is incorrectly set also clear PD_ALL_VISIBLE. offnum identifies the
+ * offending tuple for the warning, or InvalidOffsetNumber if the corruption
+ * is not specific to a tuple.
  *
  * Must be called while holding an exclusive lock on the heap buffer. Dead
  * items and not all-visible tuples must have been discovered under that same
- * lock. Although we do not hold a lock on the VM buffer, it is pinned, and
- * the heap buffer is exclusively locked, ensuring that no other backend can
- * update the VM bits corresponding to this heap page.
+ * lock. vmbuffer must be pinned and contain the VM page for buffer's block.
+ * Although we do not hold a lock on the VM buffer, the heap buffer is
+ * exclusively locked, ensuring that no other backend can update the VM bits
+ * corresponding to this heap page.
  *
  * This function makes changes to the VM and, potentially, the heap page, and
- * WAL-logs them in its own critical section.
+ * WAL-logs them in its own critical section. Callers that cache the VM status
+ * of the page must discard it after calling this.
  */
-static void
-heap_page_fix_vm_corruption(PruneState *prstate, OffsetNumber offnum,
+void
+heap_page_fix_vm_corruption(Relation relation, Buffer buffer, Buffer vmbuffer,
+							OffsetNumber offnum,
 							VMCorruptionType corruption_type)
 {
-	const char *relname = RelationGetRelationName(prstate->relation);
+	const char *relname = RelationGetRelationName(relation);
+	Page		page = BufferGetPage(buffer);
+	BlockNumber block = BufferGetBlockNumber(buffer);
 	bool		do_clear_vm = false;
 	bool		do_clear_heap = false;
 
-	Assert(BufferIsLockedByMeInMode(prstate->buffer, BUFFER_LOCK_EXCLUSIVE));
+	Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
+	Assert(visibilitymap_pin_ok(block, vmbuffer));
 
 	switch (corruption_type)
 	{
@@ -899,7 +886,7 @@ heap_page_fix_vm_corruption(PruneState *prstate, OffsetNumber offnum,
 					(errcode(ERRCODE_DATA_CORRUPTED),
 					 errmsg("dead line pointer found on page marked all-visible"),
 					 errcontext("relation \"%s\", page %u, tuple %u",
-								relname, prstate->block, offnum)));
+								relname, block, offnum)));
 			do_clear_vm = true;
 			do_clear_heap = true;
 			break;
@@ -923,7 +910,7 @@ heap_page_fix_vm_corruption(PruneState *prstate, OffsetNumber offnum,
 					(errcode(ERRCODE_DATA_CORRUPTED),
 					 errmsg("tuple not visible to all transactions found on page marked all-visible"),
 					 errcontext("relation \"%s\", page %u, tuple %u",
-								relname, prstate->block, offnum)));
+								relname, block, offnum)));
 			do_clear_vm = true;
 			do_clear_heap = true;
 			break;
@@ -938,13 +925,14 @@ heap_page_fix_vm_corruption(PruneState *prstate, OffsetNumber offnum,
 			 * that we have the buffer lock before concluding that the VM is
 			 * corrupt.
 			 */
-			Assert(!PageIsAllVisible(prstate->page));
-			Assert(prstate->old_vmbits & VISIBILITYMAP_VALID_BITS);
+			Assert(!PageIsAllVisible(page));
+			Assert(visibilitymap_get_status(relation, block, &vmbuffer) &
+				   VISIBILITYMAP_VALID_BITS);
 			ereport(WARNING,
 					(errcode(ERRCODE_DATA_CORRUPTED),
 					 errmsg("page is not marked all-visible but visibility map bit is set"),
 					 errcontext("relation \"%s\", page %u",
-								relname, prstate->block)));
+								relname, block)));
 			do_clear_vm = true;
 			break;
 
@@ -961,7 +949,7 @@ heap_page_fix_vm_corruption(PruneState *prstate, OffsetNumber offnum,
 					(errcode(ERRCODE_DATA_CORRUPTED),
 					 errmsg("page marked all-frozen in the visibility map is not all-frozen"),
 					 errcontext("relation \"%s\", page %u",
-								relname, prstate->block)));
+								relname, block)));
 			do_clear_vm = true;
 			break;
 	}
@@ -969,21 +957,20 @@ heap_page_fix_vm_corruption(PruneState *prstate, OffsetNumber offnum,
 	Assert(do_clear_heap || do_clear_vm);
 
 	if (do_clear_vm)
-		LockBuffer(prstate->vmbuffer, BUFFER_LOCK_EXCLUSIVE);
+		LockBuffer(vmbuffer, BUFFER_LOCK_EXCLUSIVE);
 
 	START_CRIT_SECTION();
 
 	if (do_clear_heap)
 	{
-		Assert(PageIsAllVisible(prstate->page));
-		PageClearAllVisible(prstate->page);
-		MarkBufferDirty(prstate->buffer);
+		Assert(PageIsAllVisible(page));
+		PageClearAllVisible(page);
+		MarkBufferDirty(buffer);
 	}
 
 	if (do_clear_vm)
 	{
-		(void) visibilitymap_clear(prstate->relation->rd_locator,
-								   prstate->block, prstate->vmbuffer,
+		(void) visibilitymap_clear(relation->rd_locator, block, vmbuffer,
 								   VISIBILITYMAP_VALID_BITS);
 
 		/*
@@ -993,7 +980,7 @@ heap_page_fix_vm_corruption(PruneState *prstate, OffsetNumber offnum,
 		 * Logging it ensures the VM is also cleared on the standby. If we log
 		 * it, we have to mark it dirty.
 		 */
-		MarkBufferDirty(prstate->vmbuffer);
+		MarkBufferDirty(vmbuffer);
 	}
 
 	/*
@@ -1004,34 +991,30 @@ heap_page_fix_vm_corruption(PruneState *prstate, OffsetNumber offnum,
 	 * images of the pages we changed. VM corruption is rare, so the extra WAL
 	 * does not matter.
 	 */
-	if (RelationNeedsWAL(prstate->relation))
+	if (RelationNeedsWAL(relation))
 	{
 		XLogRecPtr	recptr;
 		uint8		block_id = 0;
 
 		XLogBeginInsert();
 		if (do_clear_heap)
-			XLogRegisterBuffer(block_id++, prstate->buffer,
+			XLogRegisterBuffer(block_id++, buffer,
 							   REGBUF_FORCE_IMAGE | REGBUF_STANDARD);
 		if (do_clear_vm)
-			XLogRegisterBuffer(block_id++, prstate->vmbuffer,
-							   REGBUF_FORCE_IMAGE);
+			XLogRegisterBuffer(block_id++, vmbuffer, REGBUF_FORCE_IMAGE);
 
 		recptr = XLogInsert(RM_XLOG_ID, XLOG_FPI);
 
 		if (do_clear_heap)
-			PageSetLSN(prstate->page, recptr);
+			PageSetLSN(page, recptr);
 		if (do_clear_vm)
-			PageSetLSN(BufferGetPage(prstate->vmbuffer), recptr);
+			PageSetLSN(BufferGetPage(vmbuffer), recptr);
 	}
 
 	END_CRIT_SECTION();
 
 	if (do_clear_vm)
-	{
-		LockBuffer(prstate->vmbuffer, BUFFER_LOCK_UNLOCK);
-		prstate->old_vmbits = 0;
-	}
+		LockBuffer(vmbuffer, BUFFER_LOCK_UNLOCK);
 }
 
 /*
@@ -1237,8 +1220,12 @@ heap_page_prune_and_freeze(PruneFreezeParams *params,
 	 */
 	if ((prstate.old_vmbits & VISIBILITYMAP_VALID_BITS) &&
 		!PageIsAllVisible(prstate.page))
-		heap_page_fix_vm_corruption(&prstate, InvalidOffsetNumber,
+	{
+		heap_page_fix_vm_corruption(prstate.relation, prstate.buffer,
+									prstate.vmbuffer, InvalidOffsetNumber,
 									VM_CORRUPT_MISSING_PAGE_HINT);
+		prstate.old_vmbits = 0;
+	}
 
 	/*
 	 * If the page is already all-frozen, or already all-visible when freezing
@@ -1336,8 +1323,12 @@ heap_page_prune_and_freeze(PruneFreezeParams *params,
 	if (prstate.attempt_freeze && prstate.set_all_visible &&
 		!prstate.set_all_frozen &&
 		(prstate.old_vmbits & VISIBILITYMAP_ALL_FROZEN))
-		heap_page_fix_vm_corruption(&prstate, InvalidOffsetNumber,
+	{
+		heap_page_fix_vm_corruption(prstate.relation, prstate.buffer,
+									prstate.vmbuffer, InvalidOffsetNumber,
 									VM_CORRUPT_STALE_ALL_FROZEN);
+		prstate.old_vmbits = 0;
+	}
 
 	do_set_vm = heap_page_will_set_vm(&prstate, params->reason, do_prune, do_freeze);
 
@@ -1838,8 +1829,12 @@ heap_prune_record_prunable(PruneState *prstate, TransactionId xid,
 	 * prunable items.
 	 */
 	if (PageIsAllVisible(prstate->page))
-		heap_page_fix_vm_corruption(prstate, offnum,
+	{
+		heap_page_fix_vm_corruption(prstate->relation, prstate->buffer,
+									prstate->vmbuffer, offnum,
 									VM_CORRUPT_TUPLE_VISIBILITY);
+		prstate->old_vmbits = 0;
+	}
 }
 
 /* Record line pointer to be redirected */
@@ -1931,7 +1926,12 @@ heap_prune_record_dead_or_unused(PruneState *prstate, OffsetNumber offnum,
 	 * cover tuples that are directly marked LP_UNUSED via mark_unused_now.
 	 */
 	if (PageIsAllVisible(prstate->page))
-		heap_page_fix_vm_corruption(prstate, offnum, VM_CORRUPT_LPDEAD);
+	{
+		heap_page_fix_vm_corruption(prstate->relation, prstate->buffer,
+									prstate->vmbuffer, offnum,
+									VM_CORRUPT_LPDEAD);
+		prstate->old_vmbits = 0;
+	}
 }
 
 /* Record line pointer to be marked unused */
@@ -2168,7 +2168,12 @@ heap_prune_record_unchanged_lp_dead(PruneState *prstate, OffsetNumber offnum)
 	 * items.
 	 */
 	if (PageIsAllVisible(prstate->page))
-		heap_page_fix_vm_corruption(prstate, offnum, VM_CORRUPT_LPDEAD);
+	{
+		heap_page_fix_vm_corruption(prstate->relation, prstate->buffer,
+									prstate->vmbuffer, offnum,
+									VM_CORRUPT_LPDEAD);
+		prstate->old_vmbits = 0;
+	}
 }
 
 /*
diff --git a/src/include/access/heapam.h b/src/include/access/heapam.h
index 9e35961fb9e..a968b1fa78b 100644
--- a/src/include/access/heapam.h
+++ b/src/include/access/heapam.h
@@ -255,6 +255,24 @@ typedef enum
 	PRUNE_VACUUM_CLEANUP,		/* VACUUM 2nd heap pass */
 } PruneReason;
 
+/*
+ * Type of visibility map corruption detected on a heap page and its
+ * associated VM page. Passed to heap_page_fix_vm_corruption() so the caller
+ * can specify what it found rather than having the function rederive the
+ * corruption from page state.
+ */
+typedef enum VMCorruptionType
+{
+	/* VM bits are set but the heap page-level PD_ALL_VISIBLE flag is not */
+	VM_CORRUPT_MISSING_PAGE_HINT,
+	/* LP_DEAD line pointers found on a page marked all-visible */
+	VM_CORRUPT_LPDEAD,
+	/* Tuple not visible to all transactions on a page marked all-visible */
+	VM_CORRUPT_TUPLE_VISIBILITY,
+	/* Page marked all-frozen in the VM but not actually all-frozen */
+	VM_CORRUPT_STALE_ALL_FROZEN,
+} VMCorruptionType;
+
 /*
  * Input parameters to heap_page_prune_and_freeze()
  */
@@ -448,6 +466,9 @@ extern void heap_page_prune_and_freeze(PruneFreezeParams *params,
 									   OffsetNumber *off_loc,
 									   TransactionId *new_relfrozen_xid,
 									   MultiXactId *new_relmin_mxid);
+extern void heap_page_fix_vm_corruption(Relation relation, Buffer buffer,
+										Buffer vmbuffer, OffsetNumber offnum,
+										VMCorruptionType corruption_type);
 extern void heap_page_prune_execute(Buffer buffer, bool lp_truncate_only,
 									OffsetNumber *redirected, int nredirected,
 									OffsetNumber *nowdead, int ndead,
-- 
2.43.0

From f72a487d75f98d852a9da10b3fdd3db9aa2cbae4 Mon Sep 17 00:00:00 2001
From: Melanie Plageman <[email protected]>
Date: Fri, 25 Sep 2026 14:57:28 -0400
Subject: [PATCH v1 04/11] Repair visibility map corruption on empty pages

lazy_scan_new_or_empty() only checked PD_ALL_VISIBLE before setting an
empty page all-visible and all-frozen. If the VM bits were already set
while PD_ALL_VISIBLE was clear, the corruption went unreported, and
visibilitymap_set() could be a no-op, leaving the VM buffer clean when
log_heap_prune_and_freeze() expects it to be dirty.

Detect that case and repair it with heap_page_fix_vm_corruption() before
setting the page all-visible. Also set the VM when PD_ALL_VISIBLE is
already set but the VM bits are not both set.
---
 src/backend/access/heap/vacuumlazy.c | 45 ++++++++++++++++++++--------
 1 file changed, 33 insertions(+), 12 deletions(-)

diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 997d84a77b3..ad9aa64be5a 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -1927,6 +1927,10 @@ lazy_scan_new_or_empty(LVRelState *vacrel, Buffer buf, BlockNumber blkno,
 
 	if (PageIsEmpty(page))
 	{
+		uint8		old_vmbits;
+		uint8		vmflags = VISIBILITYMAP_ALL_VISIBLE |
+			VISIBILITYMAP_ALL_FROZEN;
+
 		/*
 		 * It seems likely that caller will always be able to get a cleanup
 		 * lock on an empty page.  But don't take any chances -- escalate to
@@ -1952,7 +1956,21 @@ lazy_scan_new_or_empty(LVRelState *vacrel, Buffer buf, BlockNumber blkno,
 		 * Unlike new pages, empty pages are always set all-visible and
 		 * all-frozen.
 		 */
-		if (!PageIsAllVisible(page))
+		old_vmbits = visibilitymap_get_status(vacrel->rel, blkno, &vmbuffer);
+
+		/*
+		 * If the VM is set but PD_ALL_VISIBLE is clear, fix that corruption
+		 * first, so that we set both below in a single WAL-logged operation.
+		 */
+		if ((old_vmbits & VISIBILITYMAP_VALID_BITS) && !PageIsAllVisible(page))
+		{
+			heap_page_fix_vm_corruption(vacrel->rel, buf, vmbuffer,
+										InvalidOffsetNumber,
+										VM_CORRUPT_MISSING_PAGE_HINT);
+			old_vmbits = 0;
+		}
+
+		if (!PageIsAllVisible(page) || old_vmbits != vmflags)
 		{
 			/* Lock vmbuffer before entering critical section */
 			LockBuffer(vmbuffer, BUFFER_LOCK_EXCLUSIVE);
@@ -1964,11 +1982,10 @@ lazy_scan_new_or_empty(LVRelState *vacrel, Buffer buf, BlockNumber blkno,
 
 			PageSetAllVisible(page);
 			PageClearPrunable(page);
-			(void) visibilitymap_set(blkno,
-									 vmbuffer,
-									 VISIBILITYMAP_ALL_VISIBLE |
-									 VISIBILITYMAP_ALL_FROZEN,
-									 vacrel->rel->rd_locator);
+
+			old_vmbits = visibilitymap_set(blkno, vmbuffer, vmflags,
+										   vacrel->rel->rd_locator);
+			Assert(old_vmbits != vmflags);
 
 			/*
 			 * Emit WAL for setting PD_ALL_VISIBLE on the heap page and
@@ -1976,9 +1993,7 @@ lazy_scan_new_or_empty(LVRelState *vacrel, Buffer buf, BlockNumber blkno,
 			 */
 			if (RelationNeedsWAL(vacrel->rel))
 				log_heap_prune_and_freeze(vacrel->rel, buf,
-										  vmbuffer,
-										  VISIBILITYMAP_ALL_VISIBLE |
-										  VISIBILITYMAP_ALL_FROZEN,
+										  vmbuffer, vmflags,
 										  InvalidTransactionId, /* conflict xid */
 										  false,	/* cleanup lock */
 										  PRUNE_VACUUM_SCAN,	/* reason */
@@ -1991,9 +2006,15 @@ lazy_scan_new_or_empty(LVRelState *vacrel, Buffer buf, BlockNumber blkno,
 
 			LockBuffer(vmbuffer, BUFFER_LOCK_UNLOCK);
 
-			/* Count the newly all-frozen pages for logging */
-			vacrel->new_all_visible_pages++;
-			vacrel->new_all_visible_all_frozen_pages++;
+			/* Count only the VM bits that were newly set. */
+			if (!(old_vmbits & VISIBILITYMAP_ALL_VISIBLE))
+			{
+				vacrel->new_all_visible_pages++;
+				if (!(old_vmbits & VISIBILITYMAP_ALL_FROZEN))
+					vacrel->new_all_visible_all_frozen_pages++;
+			}
+			else if (!(old_vmbits & VISIBILITYMAP_ALL_FROZEN))
+				vacrel->new_all_frozen_pages++;
 		}
 
 		freespace = PageGetHeapFreeSpace(page);
-- 
2.43.0

From 24048b2a3f01390ff0842d347f05c4b17e425fa5 Mon Sep 17 00:00:00 2001
From: Melanie Plageman <[email protected]>
Date: Wed, 23 Sep 2026 16:42:40 -0400
Subject: [PATCH v1 05/11] Register the VM buffer whenever clearing
 PD_ALL_VISIBLE

Heap WAL records clearing PD_ALL_VISIBLE (insert, multi_insert, delete,
update) only registered the VM buffer when clearing its bits changed the
VM page. If the VM was missing on the primary (e.g. historically CREATE
DATABASE STRATEGY WAL_LOG copying from a template could cause this)
skipping the VM during redo then leaves data corruption.

Always register the VM buffer when clearing PD_ALL_VISIBLE, using
REGBUF_NO_CHANGE when the bits were already clear, so redo always has
the VM block reference and clears any divergent bit. Apply the same rule
to pg_surgery's heap_force_kill(), which logs a full-page image of the VM
even when clearing its bits was a no-op on the primary. Only update the
primary VM page's LSN when its bits changed.

Tuple locking is left as-is for now as it does not have a page-level hint
to keep synchronized. XXX: is this okay?
---
 contrib/pg_surgery/heap_surgery.c | 14 ++++--
 src/backend/access/heap/heapam.c  | 76 ++++++++++++++++++++++++-------
 2 files changed, 71 insertions(+), 19 deletions(-)

diff --git a/contrib/pg_surgery/heap_surgery.c b/contrib/pg_surgery/heap_surgery.c
index 51f3f3c49eb..489de5c1a54 100644
--- a/contrib/pg_surgery/heap_surgery.c
+++ b/contrib/pg_surgery/heap_surgery.c
@@ -155,6 +155,7 @@ heap_force_common(FunctionCallInfo fcinfo, HeapTupleForceOption heap_force_opt)
 		int			i;
 		bool		did_modify_page = false;
 		bool		did_modify_vm = false;
+		bool		all_visible_cleared = false;
 
 		CHECK_FOR_INTERRUPTS();
 
@@ -277,6 +278,7 @@ heap_force_common(FunctionCallInfo fcinfo, HeapTupleForceOption heap_force_opt)
 						did_modify_vm = true;
 
 					PageClearAllVisible(page);
+					all_visible_cleared = true;
 				}
 			}
 			else
@@ -332,9 +334,15 @@ heap_force_common(FunctionCallInfo fcinfo, HeapTupleForceOption heap_force_opt)
 
 				XLogBeginInsert();
 				XLogRegisterBuffer(0, buf, REGBUF_STANDARD | REGBUF_FORCE_IMAGE);
-				/* Include the VM page if it was modified */
-				if (did_modify_vm)
-					XLogRegisterBuffer(1, vmbuf, REGBUF_FORCE_IMAGE);
+				/*
+				 * Include the VM image whenever we cleared PD_ALL_VISIBLE, even
+				 * if its bits were already clear on the primary. A standby may
+				 * still have them set and must clear them along with the heap
+				 * hint. If the VM was unchanged, leave its LSN alone below.
+				 */
+				if (all_visible_cleared)
+					XLogRegisterBuffer(1, vmbuf, REGBUF_FORCE_IMAGE |
+									   (did_modify_vm ? 0 : REGBUF_NO_CHANGE));
 				recptr = XLogInsert(RM_XLOG_ID, XLOG_FPI);
 				if (did_modify_vm)
 					PageSetLSN(BufferGetPage(vmbuf), recptr);
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 4207f0e0e08..d2797237f4c 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -64,6 +64,7 @@ static XLogRecPtr log_heap_update(Relation reln, Buffer oldbuf,
 								  Buffer vmbuffer_new, HeapTuple oldtup,
 								  HeapTuple newtup, HeapTuple old_key_tuple,
 								  bool all_visible_cleared, bool new_all_visible_cleared,
+								  bool vmbuffer_old_modified, bool vmbuffer_new_modified,
 								  bool walLogical);
 #ifdef USE_ASSERT_CHECKING
 static void check_lock_if_inplace_updateable_rel(Relation relation,
@@ -2174,8 +2175,15 @@ heap_insert(Relation relation, HeapTuple tup, CommandId cid,
 		/* filtering by origin on a row level is much more efficient */
 		XLogSetRecordFlags(XLOG_INCLUDE_ORIGIN);
 
-		if (vmbuffer_modified)
-			XLogRegisterBuffer(HEAP_INSERT_BLKREF_VM, vmbuffer, 0);
+		/*
+		 * Register the VM buffer even if its bits were already clear, so redo
+		 * clears PD_ALL_VISIBLE's VM bits. Without this, a missing VM on
+		 * master would cause data corruption on the standby when we failed to
+		 * clear the VM there.
+		 */
+		if (clear_all_visible)
+			XLogRegisterBuffer(HEAP_INSERT_BLKREF_VM, vmbuffer,
+							   vmbuffer_modified ? 0 : REGBUF_NO_CHANGE);
 
 		recptr = XLogInsert(RM_HEAP_ID, info);
 
@@ -2486,11 +2494,14 @@ heap_multi_insert(Relation relation, TupleTableSlot **slots, int ntuples,
 		{
 			PageSetAllVisible(page);
 			PageClearPrunable(page);
-			(void) visibilitymap_set(BufferGetBlockNumber(buffer),
-									 vmbuffer,
-									 VISIBILITYMAP_ALL_VISIBLE |
-									 VISIBILITYMAP_ALL_FROZEN,
-									 relation->rd_locator);
+
+			if (visibilitymap_set(BufferGetBlockNumber(buffer),
+								  vmbuffer,
+								  VISIBILITYMAP_ALL_VISIBLE |
+								  VISIBILITYMAP_ALL_FROZEN,
+								  relation->rd_locator) !=
+				(VISIBILITYMAP_ALL_VISIBLE | VISIBILITYMAP_ALL_FROZEN))
+				vmbuffer_modified = true;
 		}
 
 		/*
@@ -2613,8 +2624,16 @@ heap_multi_insert(Relation relation, TupleTableSlot **slots, int ntuples,
 			XLogRegisterData(xlrec, tupledata - scratch.data);
 			XLogRegisterBuffer(HEAP_MULTI_INSERT_BLKREF_HEAP, buffer,
 							   REGBUF_STANDARD | bufflags);
-			if (all_frozen_set || vmbuffer_modified)
-				XLogRegisterBuffer(HEAP_MULTI_INSERT_BLKREF_VM, vmbuffer, 0);
+
+			/*
+			 * If the vmbuffer wasn't modified but we are clearing all-visible
+			 * (for example because the VM was lost), register it with
+			 * REGBUF_NO_CHANGE to keep the xlog system from complaining it
+			 * must be dirty.
+			 */
+			if (all_frozen_set || clear_all_visible)
+				XLogRegisterBuffer(HEAP_MULTI_INSERT_BLKREF_VM, vmbuffer,
+								   vmbuffer_modified ? 0 : REGBUF_NO_CHANGE);
 
 			XLogRegisterBufData(HEAP_MULTI_INSERT_BLKREF_HEAP, tupledata,
 								totaldatalen);
@@ -2625,7 +2644,7 @@ heap_multi_insert(Relation relation, TupleTableSlot **slots, int ntuples,
 			recptr = XLogInsert(RM_HEAP2_ID, info);
 
 			PageSetLSN(page, recptr);
-			if (all_frozen_set || vmbuffer_modified)
+			if (vmbuffer_modified)
 			{
 				Assert(BufferIsDirty(vmbuffer));
 				PageSetLSN(BufferGetPage(vmbuffer), recptr);
@@ -3144,8 +3163,15 @@ heap_delete(Relation relation, const ItemPointerData *tid,
 		/* filtering by origin on a row level is much more efficient */
 		XLogSetRecordFlags(XLOG_INCLUDE_ORIGIN);
 
-		if (vmbuffer_modified)
-			XLogRegisterBuffer(HEAP_DELETE_BLKREF_VM, vmbuffer, 0);
+		/*
+		 * If the vmbuffer wasn't modified but we are clearing all-visible
+		 * (for example because the VM was lost), register it with
+		 * REGBUF_NO_CHANGE to keep the xlog system from complaining it must
+		 * be dirty.
+		 */
+		if (clear_all_visible)
+			XLogRegisterBuffer(HEAP_DELETE_BLKREF_VM, vmbuffer,
+							   vmbuffer_modified ? 0 : REGBUF_NO_CHANGE);
 
 		recptr = XLogInsert(RM_HEAP_ID, XLOG_HEAP_DELETE);
 
@@ -4276,14 +4302,25 @@ heap_update(Relation relation, const ItemPointerData *otid, HeapTuple newtup,
 			log_heap_new_cid(relation, heaptup);
 		}
 
+		/*
+		 * When both heap pages are all-visible and share a VM page, that page
+		 * is registered once as VM_NEW; pass the old slot as invalid to avoid
+		 * registering the same buffer twice.
+		 */
 		recptr = log_heap_update(relation, buffer,
-								 vmbuffer_modified ? vmbuffer : InvalidBuffer,
+								 (clear_all_visible &&
+								  !(clear_all_visible_new &&
+									vmbuffer == vmbuffer_new)) ?
+								 vmbuffer : InvalidBuffer,
 								 newbuf,
-								 vmbuffer_new_modified ? vmbuffer_new : InvalidBuffer,
+								 clear_all_visible_new ?
+								 vmbuffer_new : InvalidBuffer,
 								 &oldtup, heaptup,
 								 old_key_tuple,
 								 clear_all_visible,
 								 clear_all_visible_new,
+								 vmbuffer_modified,
+								 vmbuffer_new_modified,
 								 walLogical);
 		if (newbuf != buffer)
 		{
@@ -9023,6 +9060,7 @@ log_heap_update(Relation reln, Buffer oldbuf, Buffer vmbuffer_old,
 				HeapTuple oldtup, HeapTuple newtup,
 				HeapTuple old_key_tuple,
 				bool all_visible_cleared, bool new_all_visible_cleared,
+				bool vmbuffer_old_modified, bool vmbuffer_new_modified,
 				bool walLogical)
 {
 	xl_heap_update xlrec;
@@ -9235,15 +9273,21 @@ log_heap_update(Relation reln, Buffer oldbuf, Buffer vmbuffer_old,
 	 * same VM page and both their VM bits were cleared, the caller passes
 	 * only vmbuffer_new (mirroring the heap page convention where block 0 =
 	 * new is always registered).
+	 *
+	 * A buffer is registered even if its bits were already clear, so redo
+	 * clears PD_ALL_VISIBLE's VM bits (the VM can be out-of-sync across a
+	 * cluster); use REGBUF_NO_CHANGE when the page was not modified.
 	 */
 	Assert((BufferIsInvalid(vmbuffer_old) && BufferIsInvalid(vmbuffer_new)) ||
 		   (vmbuffer_old != vmbuffer_new));
 
 	if (BufferIsValid(vmbuffer_new))
-		XLogRegisterBuffer(HEAP_UPDATE_BLKREF_VM_NEW, vmbuffer_new, 0);
+		XLogRegisterBuffer(HEAP_UPDATE_BLKREF_VM_NEW, vmbuffer_new,
+						   vmbuffer_new_modified ? 0 : REGBUF_NO_CHANGE);
 
 	if (BufferIsValid(vmbuffer_old))
-		XLogRegisterBuffer(HEAP_UPDATE_BLKREF_VM_OLD, vmbuffer_old, 0);
+		XLogRegisterBuffer(HEAP_UPDATE_BLKREF_VM_OLD, vmbuffer_old,
+						   vmbuffer_old_modified ? 0 : REGBUF_NO_CHANGE);
 
 	/* filtering by origin on a row level is much more efficient */
 	XLogSetRecordFlags(XLOG_INCLUDE_ORIGIN);
-- 
2.43.0

From 91e6c2d863b46e31f1f401a65504f968e53b9846 Mon Sep 17 00:00:00 2001
From: Melanie Plageman <[email protected]>
Date: Fri, 25 Sep 2026 13:42:58 -0400
Subject: [PATCH v1 06/11] Factor out parsing of flag-list GUCs with boolean
 compatibility

log_connections is a comma-separated list of options whose flags are
ORed together, and, for backwards compatibility with its earlier boolean
form, also accepts a boolean value on its own. Move that parsing into a
generic helper so that it can be used for other similarly structured
GUCs' check hooks int he future.

No behavior change for log_connections, except that the error detail for a
boolean value given in a list no longer names the GUC ("Cannot specify
option "on" in a list with other options."); the error message itself
already names it.
---
 src/backend/tcop/backend_startup.c | 152 +++--------------------------
 src/backend/utils/misc/guc.c       | 147 ++++++++++++++++++++++++++++
 src/include/utils/guc.h            |   3 +
 3 files changed, 163 insertions(+), 139 deletions(-)

diff --git a/src/backend/tcop/backend_startup.c b/src/backend/tcop/backend_startup.c
index 912ad7dc957..dc0dc182903 100644
--- a/src/backend/tcop/backend_startup.c
+++ b/src/backend/tcop/backend_startup.c
@@ -40,7 +40,6 @@
 #include "utils/memutils.h"
 #include "utils/ps_status.h"
 #include "utils/timeout.h"
-#include "utils/varlena.h"
 
 /* GUCs */
 bool		Trace_connection_negotiation = false;
@@ -64,7 +63,6 @@ static void ProcessCancelRequestPacket(Port *port, void *pkt, int pktlen);
 static void SendNegotiateProtocolVersion(List *unrecognized_protocol_options);
 static void process_startup_packet_die(SIGNAL_ARGS);
 static void StartupPacketTimeoutHandler(void);
-static bool validate_log_connections_options(List *elemlist, uint32 *flags);
 
 /*
  * Entry point for a new backend process.
@@ -1006,150 +1004,26 @@ StartupPacketTimeoutHandler(void)
 	_exit(1);
 }
 
-/*
- * Helper for the log_connections GUC check hook.
- *
- * `elemlist` is a listified version of the string input passed to the
- * log_connections GUC check hook, check_log_connections().
- * check_log_connections() is responsible for cleaning up `elemlist`.
- *
- * validate_log_connections_options() returns false if an error was
- * encountered and the GUC input could not be validated and true otherwise.
- *
- * `flags` returns the flags that should be stored in the log_connections GUC
- * by its assign hook.
- */
-static bool
-validate_log_connections_options(List *elemlist, uint32 *flags)
-{
-	ListCell   *l;
-	char	   *item;
-
-	/*
-	 * For backwards compatibility, we accept these tokens by themselves.
-	 *
-	 * Prior to PostgreSQL 18, log_connections was a boolean GUC that accepted
-	 * any unambiguous substring of 'true', 'false', 'yes', 'no', 'on', and
-	 * 'off'. Since log_connections became a list of strings in 18, we only
-	 * accept complete option strings.
-	 */
-	static const struct config_enum_entry compat_options[] = {
-		{"off", 0},
-		{"false", 0},
-		{"no", 0},
-		{"0", 0},
-		{"on", LOG_CONNECTION_ON},
-		{"true", LOG_CONNECTION_ON},
-		{"yes", LOG_CONNECTION_ON},
-		{"1", LOG_CONNECTION_ON},
-	};
-
-	*flags = 0;
-
-	/* If an empty string was passed, we're done */
-	if (list_length(elemlist) == 0)
-		return true;
-
-	/*
-	 * Now check for the backwards compatibility options. They must always be
-	 * specified on their own, so we error out if the first option is a
-	 * backwards compatibility option and other options are also specified.
-	 */
-	item = linitial(elemlist);
-
-	for (size_t i = 0; i < lengthof(compat_options); i++)
-	{
-		struct config_enum_entry option = compat_options[i];
-
-		if (pg_strcasecmp(item, option.name) != 0)
-			continue;
-
-		if (list_length(elemlist) > 1)
-		{
-			GUC_check_errdetail("Cannot specify log_connections option \"%s\" in a list with other options.",
-								item);
-			return false;
-		}
-
-		*flags = option.val;
-		return true;
-	}
-
-	/* Now check the aspect options. The empty string was already handled */
-	foreach(l, elemlist)
-	{
-		static const struct config_enum_entry options[] = {
-			{"receipt", LOG_CONNECTION_RECEIPT},
-			{"authentication", LOG_CONNECTION_AUTHENTICATION},
-			{"authorization", LOG_CONNECTION_AUTHORIZATION},
-			{"setup_durations", LOG_CONNECTION_SETUP_DURATIONS},
-			{"all", LOG_CONNECTION_ALL},
-		};
-
-		item = lfirst(l);
-		for (size_t i = 0; i < lengthof(options); i++)
-		{
-			struct config_enum_entry option = options[i];
-
-			if (pg_strcasecmp(item, option.name) == 0)
-			{
-				*flags |= option.val;
-				goto next;
-			}
-		}
-
-		GUC_check_errdetail("Invalid option \"%s\".", item);
-		return false;
-
-next:	;
-	}
-
-	return true;
-}
-
-
 /*
  * GUC check hook for log_connections
+ *
+ * Prior to PostgreSQL 18, log_connections was a boolean GUC; a boolean value
+ * meaning true selects LOG_CONNECTION_ON.
  */
 bool
 check_log_connections(char **newval, void **extra, GucSource source)
 {
-	uint32		flags;
-	char	   *rawstring;
-	List	   *elemlist;
-	bool		success;
-
-	/* Need a modifiable copy of string */
-	rawstring = pstrdup(*newval);
-
-	if (!SplitIdentifierString(rawstring, ',', &elemlist))
-	{
-		GUC_check_errdetail("Invalid list syntax in parameter \"%s\".", "log_connections");
-		pfree(rawstring);
-		list_free(elemlist);
-		return false;
-	}
-
-	/* Validation logic is all in the helper */
-	success = validate_log_connections_options(elemlist, &flags);
-
-	/* Time for cleanup */
-	pfree(rawstring);
-	list_free(elemlist);
-
-	if (!success)
-		return false;
-
-	/*
-	 * We succeeded, so allocate `extra` and save the flags there for use by
-	 * assign_log_connections().
-	 */
-	*extra = guc_malloc(LOG, sizeof(int));
-	if (!*extra)
-		return false;
-	*((int *) *extra) = flags;
+	static const struct config_enum_entry options[] = {
+		{"receipt", LOG_CONNECTION_RECEIPT},
+		{"authentication", LOG_CONNECTION_AUTHENTICATION},
+		{"authorization", LOG_CONNECTION_AUTHORIZATION},
+		{"setup_durations", LOG_CONNECTION_SETUP_DURATIONS},
+		{"all", LOG_CONNECTION_ALL},
+		{NULL, 0}
+	};
 
-	return true;
+	return check_flag_list_guc(newval, extra, "log_connections", options,
+							   true, LOG_CONNECTION_ON);
 }
 
 /*
diff --git a/src/backend/utils/misc/guc.c b/src/backend/utils/misc/guc.c
index 21c9ccec1e2..c32929c93eb 100644
--- a/src/backend/utils/misc/guc.c
+++ b/src/backend/utils/misc/guc.c
@@ -51,6 +51,7 @@
 #include "utils/guc_tables.h"
 #include "utils/memutils.h"
 #include "utils/timestamp.h"
+#include "utils/varlena.h"
 
 
 #define CONFIG_FILENAME "postgresql.conf"
@@ -2961,6 +2962,152 @@ config_enum_lookup_by_name(const struct config_enum *record, const char *value,
 	return false;
 }
 
+/*
+ * Helper for check_flag_list_guc(): compute the flags selected by 'elemlist',
+ * the listified value of the GUC. Returns false, with the GUC error detail
+ * set, if the list is invalid.
+ */
+static bool
+validate_flag_list_guc_options(List *elemlist,
+							   const struct config_enum_entry *options,
+							   bool boolean_compat, int on_value, int *flags)
+{
+	ListCell   *l;
+	char	   *item;
+
+	/*
+	 * For backwards compatibility with GUCs that used to be booleans, we
+	 * accept these tokens by themselves. A boolean GUC accepts any
+	 * unambiguous substring of 'true', 'false', 'yes', 'no', 'on', and 'off',
+	 * but here we only accept complete option strings.
+	 */
+	static const struct config_enum_entry compat_options[] = {
+		{"off", false},
+		{"false", false},
+		{"no", false},
+		{"0", false},
+		{"on", true},
+		{"true", true},
+		{"yes", true},
+		{"1", true},
+	};
+
+	*flags = 0;
+
+	/* If an empty string was passed, we're done */
+	if (list_length(elemlist) == 0)
+		return true;
+
+	/*
+	 * If the GUC used to be a boolean, check for the backwards compatibility
+	 * options. They must always be specified on their own, so we error out if
+	 * the first option is a backwards compatibility option and other options
+	 * are also specified.
+	 */
+	if (boolean_compat)
+	{
+		item = linitial(elemlist);
+
+		for (size_t i = 0; i < lengthof(compat_options); i++)
+		{
+			if (pg_strcasecmp(item, compat_options[i].name) != 0)
+				continue;
+
+			if (list_length(elemlist) > 1)
+			{
+				GUC_check_errdetail("Cannot specify option \"%s\" in a list with other options.",
+									item);
+				return false;
+			}
+
+			*flags = compat_options[i].val ? on_value : 0;
+			return true;
+		}
+	}
+
+	/* Now check the regular options. The empty string was already handled */
+	foreach(l, elemlist)
+	{
+		const struct config_enum_entry *option;
+
+		item = lfirst(l);
+		for (option = options; option->name; option++)
+		{
+			if (pg_strcasecmp(item, option->name) == 0)
+				break;
+		}
+
+		if (!option->name)
+		{
+			GUC_check_errdetail("Invalid option \"%s\".", item);
+			return false;
+		}
+
+		*flags |= option->val;
+	}
+
+	return true;
+}
+
+/*
+ * Check hook body for a list-valued GUC whose items are flags.
+ *
+ * *newval is a comma-separated list of option names from 'options', an array
+ * terminated by an entry with a NULL name. The flags of the listed options
+ * are ORed together and stored in *extra, as an int, for the GUC's assign
+ * hook.
+ *
+ * If 'boolean_compat' is true, for backwards compatibility with a GUC that
+ * used to be a boolean, a boolean value ('on', 'true', 'yes', '1', or their
+ * negations) is also accepted on its own, selecting 'on_value' or no flags
+ * respectively. Otherwise 'on_value' is ignored.
+ *
+ * 'name' is the GUC's name, for error messages.
+ */
+bool
+check_flag_list_guc(char **newval, void **extra, const char *name,
+					const struct config_enum_entry *options,
+					bool boolean_compat, int on_value)
+{
+	int			flags;
+	char	   *rawstring;
+	List	   *elemlist;
+	bool		success;
+
+	/* Need a modifiable copy of string */
+	rawstring = pstrdup(*newval);
+
+	if (!SplitIdentifierString(rawstring, ',', &elemlist))
+	{
+		GUC_check_errdetail("Invalid list syntax in parameter \"%s\".", name);
+		pfree(rawstring);
+		list_free(elemlist);
+		return false;
+	}
+
+	/* Validation logic is all in the helper */
+	success = validate_flag_list_guc_options(elemlist, options,
+											 boolean_compat, on_value, &flags);
+
+	/* Time for cleanup */
+	pfree(rawstring);
+	list_free(elemlist);
+
+	if (!success)
+		return false;
+
+	/*
+	 * We succeeded, so allocate `extra` and save the flags there for use by
+	 * the assign hook.
+	 */
+	*extra = guc_malloc(LOG, sizeof(int));
+	if (!*extra)
+		return false;
+	*((int *) *extra) = flags;
+
+	return true;
+}
+
 
 /*
  * Return a palloc'd string listing all the available options for an enum GUC
diff --git a/src/include/utils/guc.h b/src/include/utils/guc.h
index 164efba6b51..6401dfcf030 100644
--- a/src/include/utils/guc.h
+++ b/src/include/utils/guc.h
@@ -444,6 +444,9 @@ extern bool parse_int(const char *value, int *result, int flags,
 					  const char **hintmsg);
 extern bool parse_real(const char *value, double *result, int flags,
 					   const char **hintmsg);
+extern bool check_flag_list_guc(char **newval, void **extra, const char *name,
+								const struct config_enum_entry *options,
+								bool boolean_compat, int on_value);
 extern int	set_config_option(const char *name, const char *value,
 							  GucContext context, GucSource source,
 							  GucAction action, bool changeVal, int elevel,
-- 
2.43.0

From cee8b66b68f68bb2acc63185031463dd4f570b09 Mon Sep 17 00:00:00 2001
From: Melanie Plageman <[email protected]>
Date: Thu, 24 Sep 2026 15:59:55 -0400
Subject: [PATCH v1 07/11] Make zero_damaged_pages selectable per relation fork

zero_damaged_pages was a boolean that, when on, zeroed a damaged page of
any relation fork. Recovering from a torn or corrupt auxiliary page (for
example a visibility map page) therefore required disarming the protection
for the table's main data as well.

Turn it into a comma-separated list of forks -- "main", "fsm", "vm",
"init", or "all" -- for which a damaged page header is zeroed (with a
warning) instead of raising an error. For backward compatibility a
boolean value is still accepted on its own: "on" selects every fork and
"off" none, and existing settings keep working.

This lets an operator recover from, say, a corrupt visibility map page
with zero_damaged_pages = 'vm' while a corrupt heap page still errors
out. This will make it more defensible to stop reading the VM with
RBM_ZERO_ON_ERROR in redo.
---
 doc/src/sgml/config.sgml                  | 35 ++++++++++----
 src/backend/storage/buffer/bufmgr.c       | 57 ++++++++++++++++++++++-
 src/backend/storage/smgr/md.c             |  2 +-
 src/backend/utils/misc/guc_parameters.dat | 14 +++---
 src/include/storage/bufmgr.h              |  3 +-
 src/include/utils/guc_hooks.h             |  3 ++
 src/test/modules/test_aio/test_aio.c      |  2 +-
 7 files changed, 96 insertions(+), 20 deletions(-)

diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index 0165eb9ec02..36edb226ccb 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -13333,7 +13333,7 @@ LOG:  CleanUpLock: deleting: lock(0xb7acd844) id(24688,24696,0,0,0,1)
      </varlistentry>
 
     <varlistentry id="guc-zero-damaged-pages" xreflabel="zero_damaged_pages">
-      <term><varname>zero_damaged_pages</varname> (<type>boolean</type>)
+      <term><varname>zero_damaged_pages</varname> (<type>string</type>)
       <indexterm>
        <primary><varname>zero_damaged_pages</varname> configuration parameter</primary>
       </indexterm>
@@ -13342,18 +13342,35 @@ LOG:  CleanUpLock: deleting: lock(0xb7acd844) id(24688,24696,0,0,0,1)
        <para>
         Detection of a damaged page header normally causes
         <productname>PostgreSQL</productname> to report an error, aborting the current
-        transaction.  Setting <varname>zero_damaged_pages</varname> to on causes
-        the system to instead report a warning, zero out the damaged
-        page in memory, and continue processing.  This behavior <emphasis>will destroy data</emphasis>,
-        namely all the rows on the damaged page.  However, it does allow you to get
+        transaction.  This parameter is a comma-separated list of the relation
+        forks for which the system instead reports a warning, zeroes out the
+        damaged page in memory, and continues processing.  The recognized fork
+        names are <literal>main</literal> (the table or index data),
+        <literal>fsm</literal> (the free space map), <literal>vm</literal>
+        (the visibility map), and <literal>init</literal> (the init fork).
+        <literal>all</literal> selects every fork.  The server always zeroes
+        damaged free space map pages regardless of this setting, and it does
+        not normally read init fork pages through shared buffers, so
+        <literal>fsm</literal> and <literal>init</literal> only affect
+        functions that read a specific fork directly, such as
+        <xref linkend="pageinspect"/>'s <function>get_raw_page()</function>.
+        For backward compatibility, <literal>on</literal> (every fork) and
+        <literal>off</literal> (no fork, the default) are also accepted, but
+        only on their own.
+       </para>
+       <para>
+        Zeroing a page <emphasis>will destroy data</emphasis>, namely all the
+        rows on the damaged page.  However, it does allow you to get
         past the error and retrieve rows from any undamaged pages that might
         be present in the table.  It is useful for recovering data if
         corruption has occurred due to a hardware or software error.  You should
-        generally not set this on until you have given up hope of recovering
-        data from the damaged pages of a table.  Zeroed-out pages are not
+        generally not enable this for a fork until you have given up hope of
+        recovering data from the damaged pages of a table.  Scoping the setting
+        to a single fork, such as <literal>vm</literal>, lets you recover from a
+        damaged auxiliary page without also disarming this protection for the
+        table's main data.  Zeroed-out pages are not
         forced to disk so it is recommended to recreate the table or
-        the index before turning this parameter off again.  The
-        default setting is <literal>off</literal>.
+        the index before turning this parameter off again.
         Only superusers and users with the appropriate <literal>SET</literal>
         privilege can change this setting.
        </para>
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 5c82865a084..27539b564f0 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -46,6 +46,7 @@
 #include "catalog/storage.h"
 #include "catalog/storage_xlog.h"
 #include "common/hashfn.h"
+#include "common/relpath.h"
 #include "executor/instrument.h"
 #include "lib/binaryheap.h"
 #include "miscadmin.h"
@@ -64,6 +65,8 @@
 #include "storage/read_stream.h"
 #include "storage/smgr.h"
 #include "storage/standby.h"
+#include "utils/guc.h"
+#include "utils/guc_hooks.h"
 #include "utils/memdebug.h"
 #include "utils/ps_status.h"
 #include "utils/rel.h"
@@ -186,7 +189,15 @@ typedef struct SMgrSortArray
 } SMgrSortArray;
 
 /* GUC variables */
-bool		zero_damaged_pages = false;
+
+/*
+ * zero_damaged_pages is a list of relation forks for which a damaged page
+ * header is zeroed (with a warning) instead of raising an error. The raw
+ * GUC string is parsed into zero_damaged_pages_forks, a bitmask indexed by
+ * ForkNumber (1 << forknum).
+ */
+char	   *zero_damaged_pages_string;
+int			zero_damaged_pages_forks = 0;
 int			bgwriter_lru_maxpages = 100;
 double		bgwriter_lru_multiplier = 2.0;
 bool		track_io_timing = false;
@@ -1987,7 +1998,7 @@ AsyncReadBuffers(ReadBuffersOperation *operation, int *nblocks_progress)
 	 * zero_damaged_pages, so we can report different log levels / error codes
 	 * for zero_damaged_pages and ZERO_ON_ERROR.
 	 */
-	if (zero_damaged_pages)
+	if (zero_damaged_pages_forks & (1 << forknum))
 		flags |= READ_BUFFERS_ZERO_ON_ERROR;
 
 	/*
@@ -9005,3 +9016,45 @@ const PgAioHandleCallbacks aio_local_buffer_readv_cb = {
 	.complete_local = local_buffer_readv_complete,
 	.report = buffer_readv_report,
 };
+
+/* Fork bitmask selecting every relation fork */
+#define ZERO_DAMAGED_PAGES_ALL_FORKS ((1 << (MAX_FORKNUM + 1)) - 1)
+
+/*
+ * GUC check_hook for zero_damaged_pages.
+ *
+ * The value is a comma-separated list of relation fork names ("main", "fsm",
+ * "vm", "init"), or "all", for which a damaged page header is zeroed instead
+ * of raising an error. The resulting fork bitmask is stashed in *extra for
+ * the assign hook.
+ *
+ * Prior to PostgreSQL 20, zero_damaged_pages was a boolean GUC; a boolean
+ * value meaning true selects every fork.
+ */
+bool
+check_zero_damaged_pages(char **newval, void **extra, GucSource source)
+{
+	static const struct config_enum_entry options[] = {
+		{"main", 1 << MAIN_FORKNUM},
+		{"fsm", 1 << FSM_FORKNUM},
+		{"vm", 1 << VISIBILITYMAP_FORKNUM},
+		{"init", 1 << INIT_FORKNUM},
+		{"all", ZERO_DAMAGED_PAGES_ALL_FORKS},
+		{NULL, 0}
+	};
+
+	StaticAssertDecl(lengthof(options) == MAX_FORKNUM + 3,
+					 "zero_damaged_pages must accept every fork name");
+
+	return check_flag_list_guc(newval, extra, "zero_damaged_pages", options,
+							   true, ZERO_DAMAGED_PAGES_ALL_FORKS);
+}
+
+/*
+ * GUC assign_hook for zero_damaged_pages.
+ */
+void
+assign_zero_damaged_pages(const char *newval, void *extra)
+{
+	zero_damaged_pages_forks = *((int *) extra);
+}
diff --git a/src/backend/storage/smgr/md.c b/src/backend/storage/smgr/md.c
index 780c88c0630..9e3f15c347f 100644
--- a/src/backend/storage/smgr/md.c
+++ b/src/backend/storage/smgr/md.c
@@ -951,7 +951,7 @@ mdreadv(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
 				 * continuing to work in production builds). Afterwards we
 				 * plan to remove this code entirely.
 				 */
-				if (zero_damaged_pages || InRecovery)
+				if ((zero_damaged_pages_forks & (1 << forknum)) || InRecovery)
 				{
 					Assert(false);	/* see comment above */
 
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index c57441f7d98..22906b020ba 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -3669,12 +3669,14 @@
   options => 'xmloption_options',
 },
 
-{ name => 'zero_damaged_pages', type => 'bool', context => 'PGC_SUSET', group => 'DEVELOPER_OPTIONS',
-  short_desc => 'Continues processing past damaged page headers.',
-  long_desc => 'Detection of a damaged page header normally causes PostgreSQL to report an error, aborting the current transaction. Setting "zero_damaged_pages" to true causes the system to instead report a warning, zero out the damaged page, and continue processing. This behavior will destroy data, namely all the rows on the damaged page.',
-  flags => 'GUC_NOT_IN_SAMPLE',
-  variable => 'zero_damaged_pages',
-  boot_val => 'false',
+{ name => 'zero_damaged_pages', type => 'string', context => 'PGC_SUSET', group => 'DEVELOPER_OPTIONS',
+  short_desc => 'Continues processing past damaged page headers for the listed relation forks.',
+  long_desc => 'Detection of a damaged page header normally causes PostgreSQL to report an error, aborting the current transaction. This is a comma-separated list of relation forks ("main", "fsm", "vm", "init") for which the system instead reports a warning, zeroes out the damaged page, and continues processing; this destroys all the rows on the damaged page. "all" selects every fork. For backward compatibility, "on" (every fork) and "off" (no fork) are also accepted, but only on their own.',
+  flags => 'GUC_LIST_INPUT | GUC_NOT_IN_SAMPLE',
+  variable => 'zero_damaged_pages_string',
+  boot_val => '""',
+  check_hook => 'check_zero_damaged_pages',
+  assign_hook => 'assign_zero_damaged_pages',
 },
 
 ]
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 6837b35fc6d..7d6106ffd7e 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -162,7 +162,8 @@ typedef struct WritebackContext WritebackContext;
 extern PGDLLIMPORT int NBuffers;
 
 /* in bufmgr.c */
-extern PGDLLIMPORT bool zero_damaged_pages;
+extern PGDLLIMPORT char *zero_damaged_pages_string;
+extern PGDLLIMPORT int zero_damaged_pages_forks;
 extern PGDLLIMPORT int bgwriter_lru_maxpages;
 extern PGDLLIMPORT double bgwriter_lru_multiplier;
 extern PGDLLIMPORT bool track_io_timing;
diff --git a/src/include/utils/guc_hooks.h b/src/include/utils/guc_hooks.h
index 06453a18c03..5677df7dcf9 100644
--- a/src/include/utils/guc_hooks.h
+++ b/src/include/utils/guc_hooks.h
@@ -177,5 +177,8 @@ extern bool check_synchronized_standby_slots(char **newval, void **extra,
 extern void assign_synchronized_standby_slots(const char *newval, void *extra);
 extern bool check_log_min_messages(char **newval, void **extra, GucSource source);
 extern void assign_log_min_messages(const char *newval, void *extra);
+extern bool check_zero_damaged_pages(char **newval, void **extra,
+									 GucSource source);
+extern void assign_zero_damaged_pages(const char *newval, void *extra);
 
 #endif							/* GUC_HOOKS_H */
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index 39d857557cf..083c25a9798 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -434,7 +434,7 @@ read_rel_block_ll(PG_FUNCTION_ARGS)
 
 	pgaio_io_set_handle_data_32(ioh, (uint32 *) bufs, nblocks);
 
-	if (zero_on_error | zero_damaged_pages)
+	if (zero_on_error || (zero_damaged_pages_forks & (1 << MAIN_FORKNUM)))
 		srb_flags |= READ_BUFFERS_ZERO_ON_ERROR;
 	if (ignore_checksum_failure)
 		srb_flags |= READ_BUFFERS_IGNORE_CHECKSUM_FAILURES;
-- 
2.43.0

From 4290894e69504b7f8543ed1ed0a302602b2628d9 Mon Sep 17 00:00:00 2001
From: Melanie Plageman <[email protected]>
Date: Wed, 23 Sep 2026 16:42:27 -0400
Subject: [PATCH v1 08/11] Add RBM_ZERO_ON_MISSING and read the visibility map
 with it

Add a buffer read mode, RBM_ZERO_ON_MISSING: like RBM_NORMAL for a page that
exists (a corrupt page errors), but if the block is past the fork's EOF the
fork is extended with zeroed pages during recovery rather than failing.

Use this mode when reading the VM in normal operation and recovery. The
preceding changes register VM blocks for DML clears even when they are
no-ops on the primary, and WAL-log corruption repairs. Users can select
zero_damaged_pages = 'vm' to zero a damaged VM page without also allowing
damaged main-fork pages to be zeroed.

XXX: VM truncation doesn't always register every VM block it modifies.
Look into this more.
---
 src/backend/access/heap/heapam_xlog.c   | 39 +++++++++++++++++++------
 src/backend/access/heap/visibilitymap.c | 14 +++++----
 src/backend/storage/buffer/bufmgr.c     |  3 +-
 src/include/storage/bufmgr.h            |  5 ++++
 4 files changed, 46 insertions(+), 15 deletions(-)

diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index 5fa1de09cfb..b0ca1acf2b8 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -53,11 +53,18 @@ heap_xlog_vm_clear(XLogReaderState *record,
 	 * read it. These will either apply an FPI or indicate that we should
 	 * clear the requested bits ourselves.
 	 */
-	if (XLogReadBufferForRedo(record, wal_vm_block_id,
-							  &vmbuffer) == BLK_NEEDS_REDO)
+	if (XLogReadBufferForRedoExtended(record, wal_vm_block_id,
+									  RBM_ZERO_ON_MISSING, false,
+									  &vmbuffer) == BLK_NEEDS_REDO)
 	{
+		Page		vmpage = BufferGetPage(vmbuffer);
+
+		/* initialize the page if it was extended as zeros */
+		if (PageIsNew(vmpage))
+			PageInit(vmpage, BLCKSZ, 0);
+
 		if (visibilitymap_clear(target_locator, heap_blkno, vmbuffer, flags))
-			PageSetLSN(BufferGetPage(vmbuffer), lsn);
+			PageSetLSN(vmpage, lsn);
 	}
 	if (BufferIsValid(vmbuffer))
 		UnlockReleaseBuffer(vmbuffer);
@@ -273,7 +280,7 @@ heap_xlog_prune_freeze(XLogReaderState *record)
 	 */
 	if ((vmflags & VISIBILITYMAP_VALID_BITS) &&
 		XLogReadBufferForRedoExtended(record, 1,
-									  RBM_ZERO_ON_ERROR,
+									  RBM_ZERO_ON_MISSING,
 									  false,
 									  &vmbuffer) == BLK_NEEDS_REDO)
 	{
@@ -727,7 +734,7 @@ heap_xlog_multi_insert(XLogReaderState *record)
 	 * heap_xlog_prune_freeze()).
 	 */
 	if ((xlrec->flags & XLH_INSERT_ALL_FROZEN_SET) &&
-		XLogReadBufferForRedoExtended(record, 1, RBM_ZERO_ON_ERROR, false,
+		XLogReadBufferForRedoExtended(record, 1, RBM_ZERO_ON_MISSING, false,
 									  &vmbuffer) == BLK_NEEDS_REDO)
 	{
 		Page		vmpage = BufferGetPage(vmbuffer);
@@ -812,9 +819,16 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 
 		Assert(xlrec->flags & XLH_UPDATE_NEW_ALL_VISIBLE_CLEARED);
 
-		if (XLogReadBufferForRedo(record, HEAP_UPDATE_BLKREF_VM_NEW,
-								  &vmbuffer_new) == BLK_NEEDS_REDO)
+		if (XLogReadBufferForRedoExtended(record, HEAP_UPDATE_BLKREF_VM_NEW,
+										  RBM_ZERO_ON_MISSING, false,
+										  &vmbuffer_new) == BLK_NEEDS_REDO)
 		{
+			Page		vmpage = BufferGetPage(vmbuffer_new);
+
+			/* initialize the page if it was extended as zeros */
+			if (PageIsNew(vmpage))
+				PageInit(vmpage, BLCKSZ, 0);
+
 			/*
 			 * If both the old and new heap pages were all-visible and their
 			 * VM bits are on the same VM page, that single VM page is
@@ -850,9 +864,16 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 
 		Assert(xlrec->flags & XLH_UPDATE_OLD_ALL_VISIBLE_CLEARED);
 
-		if (XLogReadBufferForRedo(record, HEAP_UPDATE_BLKREF_VM_OLD,
-								  &vmbuffer_old) == BLK_NEEDS_REDO)
+		if (XLogReadBufferForRedoExtended(record, HEAP_UPDATE_BLKREF_VM_OLD,
+										  RBM_ZERO_ON_MISSING, false,
+										  &vmbuffer_old) == BLK_NEEDS_REDO)
 		{
+			Page		vmpage = BufferGetPage(vmbuffer_old);
+
+			/* initialize the page if it was extended as zeros */
+			if (PageIsNew(vmpage))
+				PageInit(vmpage, BLCKSZ, 0);
+
 			if (visibilitymap_clear(rlocator, oldblk, vmbuffer_old,
 									VISIBILITYMAP_VALID_BITS))
 				PageSetLSN(BufferGetPage(vmbuffer_old), lsn);
diff --git a/src/backend/access/heap/visibilitymap.c b/src/backend/access/heap/visibilitymap.c
index fe5ce437e1b..73ca76462a9 100644
--- a/src/backend/access/heap/visibilitymap.c
+++ b/src/backend/access/heap/visibilitymap.c
@@ -592,9 +592,13 @@ vm_readbuf(Relation rel, BlockNumber blkno, bool extend)
 	}
 
 	/*
-	 * For reading we use ZERO_ON_ERROR mode, and initialize the page if
-	 * necessary. It's always safe to clear bits, so it's better to clear
-	 * corrupt pages than error out.
+	 * For reading we use ZERO_ON_MISSING mode, and initialize the page if
+	 * necessary. A block past the fork's end never reaches the read below: it
+	 * is either extended by vm_extend() or reported as missing. A corrupt
+	 * existing page throws an error rather than being automatically zeroed.
+	 * zero_damaged_pages can override this for the VM fork. DML changes and
+	 * corruption repairs register the VM in WAL, but truncation still has a
+	 * separate tail-clear path; see visibilitymap_prepare_truncate().
 	 *
 	 * We use the same path below to initialize pages when extending the
 	 * relation, as a concurrent extension can end up with vm_extend()
@@ -609,7 +613,7 @@ vm_readbuf(Relation rel, BlockNumber blkno, bool extend)
 	}
 	else
 		buf = ReadBufferExtended(rel, VISIBILITYMAP_FORKNUM, blkno,
-								 RBM_ZERO_ON_ERROR, NULL);
+								 RBM_ZERO_ON_MISSING, NULL);
 
 	/*
 	 * Initializing the page when needed is trickier than it looks, because of
@@ -649,7 +653,7 @@ vm_extend(Relation rel, BlockNumber vm_nblocks)
 							  EB_CREATE_FORK_IF_NEEDED |
 							  EB_CLEAR_SIZE_CACHE,
 							  vm_nblocks,
-							  RBM_ZERO_ON_ERROR);
+							  RBM_ZERO_ON_MISSING);
 
 	/*
 	 * Send a shared-inval message to force other backends to close any smgr
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 27539b564f0..4609b1b86b9 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -928,7 +928,8 @@ ReadBuffer(Relation reln, BlockNumber blockNum)
  * RBM_ZERO_AND_CLEANUP_LOCK is the same as RBM_ZERO_AND_LOCK, but acquires
  * a cleanup-strength lock on the page.
  *
- * RBM_NORMAL_NO_LOG mode is treated the same as RBM_NORMAL here.
+ * RBM_NORMAL_NO_LOG and RBM_ZERO_ON_MISSING modes are treated the same as
+ * RBM_NORMAL here.
  *
  * If strategy is not NULL, a nondefault buffer access strategy is used.
  * See buffer/README for details.
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 7d6106ffd7e..39c27c4d811 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -51,6 +51,11 @@ typedef enum
 	RBM_ZERO_ON_ERROR,			/* Read, but return an all-zeros page on error */
 	RBM_NORMAL_NO_LOG,			/* Don't log page as invalid during WAL
 								 * replay; otherwise same as RBM_NORMAL */
+	RBM_ZERO_ON_MISSING,		/* During WAL replay, extend the fork with
+								 * zeroed pages if the block doesn't exist and
+								 * accept an all-zeroes page, rather than
+								 * treating either as invalid; otherwise same
+								 * as RBM_NORMAL */
 } ReadBufferMode;
 
 /*
-- 
2.43.0

From aab002af3664640de200235ac42e11118a3dfb65 Mon Sep 17 00:00:00 2001
From: Melanie Plageman <[email protected]>
Date: Fri, 25 Sep 2026 10:33:51 -0400
Subject: [PATCH v1 09/11] Return the prior visibility map bits from
 visibilitymap_clear()

Change visibilitymap_clear() to return the visibility map bits that were set
before the clear, instead of a bool indicating whether anything was cleared.
Callers can then tell which bits were set.
---
 src/backend/access/heap/heapam.c        |  9 ++++++---
 src/backend/access/heap/heapam_xlog.c   |  3 ++-
 src/backend/access/heap/visibilitymap.c | 15 ++++++++-------
 src/include/access/visibilitymap.h      |  4 ++--
 4 files changed, 18 insertions(+), 13 deletions(-)

diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index d2797237f4c..bbee995b49b 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -3939,7 +3939,8 @@ heap_update(Relation relation, const ItemPointerData *otid, HeapTuple newtup,
 		{
 			/* It's possible all-frozen was already clear */
 			if (visibilitymap_clear(relation->rd_locator, block, vmbuffer,
-									VISIBILITYMAP_ALL_FROZEN))
+									VISIBILITYMAP_ALL_FROZEN) &
+				VISIBILITYMAP_ALL_FROZEN)
 				cleared_all_frozen = true;
 		}
 
@@ -5400,7 +5401,8 @@ heap_lock_tuple(Relation relation, HeapTuple tuple,
 	if (PageIsAllVisible(page))
 	{
 		if (visibilitymap_clear(relation->rd_locator, block, vmbuffer,
-								VISIBILITYMAP_ALL_FROZEN))
+								VISIBILITYMAP_ALL_FROZEN) &
+			VISIBILITYMAP_ALL_FROZEN)
 			cleared_all_frozen = true;
 	}
 
@@ -6193,7 +6195,8 @@ heap_lock_updated_tuple_rec(Relation rel, TransactionId priorXmax,
 		{
 			/* It's possible all-frozen was already clear */
 			if (visibilitymap_clear(rel->rd_locator, block, vmbuffer,
-									VISIBILITYMAP_ALL_FROZEN))
+									VISIBILITYMAP_ALL_FROZEN) &
+				VISIBILITYMAP_ALL_FROZEN)
 				cleared_all_frozen = true;
 		}
 
diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index b0ca1acf2b8..a50463cbaf8 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -63,7 +63,8 @@ heap_xlog_vm_clear(XLogReaderState *record,
 		if (PageIsNew(vmpage))
 			PageInit(vmpage, BLCKSZ, 0);
 
-		if (visibilitymap_clear(target_locator, heap_blkno, vmbuffer, flags))
+		if (visibilitymap_clear(target_locator, heap_blkno, vmbuffer,
+								flags) & flags)
 			PageSetLSN(vmpage, lsn);
 	}
 	if (BufferIsValid(vmbuffer))
diff --git a/src/backend/access/heap/visibilitymap.c b/src/backend/access/heap/visibilitymap.c
index 73ca76462a9..39b356512be 100644
--- a/src/backend/access/heap/visibilitymap.c
+++ b/src/backend/access/heap/visibilitymap.c
@@ -145,10 +145,11 @@ static Buffer vm_extend(Relation rel, BlockNumber vm_nblocks);
  * You must pass a buffer containing the correct map page to this function,
  * which already needs to be pinned and locked exclusively.
  *
- * This function doesn't do any I/O. Returns true if any bits have been
- * cleared and false otherwise.
+ * This function doesn't do any I/O. Returns the visibility map bits that were
+ * set for the page before this call; the bits requested in 'flags' are now
+ * cleared.
  */
-bool
+uint8
 visibilitymap_clear(RelFileLocator rlocator, BlockNumber heapBlk,
 					Buffer vmbuf, uint8 flags)
 {
@@ -158,7 +159,7 @@ visibilitymap_clear(RelFileLocator rlocator, BlockNumber heapBlk,
 	uint8		mask = flags << mapOffset;
 	Page		page;
 	char	   *map;
-	bool		cleared = false;
+	uint8		status;
 
 	/* Must never clear all_visible bit while leaving all_frozen bit set */
 	Assert(flags & VISIBILITYMAP_VALID_BITS);
@@ -178,15 +179,15 @@ visibilitymap_clear(RelFileLocator rlocator, BlockNumber heapBlk,
 	page = BufferGetPage(vmbuf);
 	map = PageGetContents(page);
 
-	if (map[mapByte] & mask)
+	status = (map[mapByte] >> mapOffset) & VISIBILITYMAP_VALID_BITS;
+	if (status & flags)
 	{
 		map[mapByte] &= ~mask;
 
 		MarkBufferDirty(vmbuf);
-		cleared = true;
 	}
 
-	return cleared;
+	return status;
 }
 
 /*
diff --git a/src/include/access/visibilitymap.h b/src/include/access/visibilitymap.h
index 165efd1c00e..f7e5b174525 100644
--- a/src/include/access/visibilitymap.h
+++ b/src/include/access/visibilitymap.h
@@ -26,8 +26,8 @@
 #define VM_ALL_FROZEN(r, b, v) \
 	((visibilitymap_get_status((r), (b), (v)) & VISIBILITYMAP_ALL_FROZEN) != 0)
 
-extern bool visibilitymap_clear(RelFileLocator rlocator, BlockNumber heapBlk,
-								Buffer vmbuf, uint8 flags);
+extern uint8 visibilitymap_clear(RelFileLocator rlocator, BlockNumber heapBlk,
+								 Buffer vmbuf, uint8 flags);
 extern void visibilitymap_pin(Relation rel, BlockNumber heapBlk,
 							  Buffer *vmbuf);
 extern bool visibilitymap_pin_ok(BlockNumber heapBlk, Buffer vmbuf);
-- 
2.43.0

From 1e71b1ea0973164761b448bb1834bdfe43208135 Mon Sep 17 00:00:00 2001
From: Melanie Plageman <[email protected]>
Date: Wed, 23 Sep 2026 16:43:20 -0400
Subject: [PATCH v1 10/11] Warn when heap redo finds a diverged visibility map

A standby's VM can diverge from the primary's (for example, via CREATE
DATABASE STRATEGY WAL_LOG), and redo silently repairs the divergence the
next time it sets or clears the bits for that page. It's helpful to
log it first. This only logs genuine data corruption on the standby that
is different than on the primary. Data corruption that is fixed on the
primary is logged there.

This only works for redo, as we don't read the existing contents of the
page when restoring an FPI.
---
 src/backend/access/heap/heapam_xlog.c | 165 ++++++++++++++++++++++----
 1 file changed, 142 insertions(+), 23 deletions(-)

diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index a50463cbaf8..eb224363168 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -18,10 +18,46 @@
 #include "access/heapam.h"
 #include "access/visibilitymap.h"
 #include "access/xlog.h"
+#include "access/xlogrecovery.h"
 #include "access/xlogutils.h"
+#include "common/relpath.h"
 #include "storage/freespace.h"
 #include "storage/standby.h"
 
+/*
+ * Report a standby whose visibility map has the all-visible bit set over a
+ * heap page that is not marked all-visible. This can happen on the standby
+ * without having happened on the primary if the VM diverges across the
+ * cluster (e.g. because of VM truncation or historical CREATE DATABASE
+ * STRATEGY WAL_LOG bugs). This is the redo counterpart of the check in
+ * heap_page_fix_vm_corruption() that runs on the primary during VACUUM.
+ *
+ * heap_all_visible is the heap page's PD_ALL_VISIBLE as it stood before this
+ * record modified it; vm_oldbits is the VM status before this record touched
+ * it; lsn is the record's LSN; and action describes what redo is doing (e.g.
+ * "clearing the visibility map bits"), all reported in the errdetail.
+ */
+static void
+heap_xlog_warn_vm_corruption(bool heap_all_visible, uint8 vm_oldbits,
+							 XLogRecPtr lsn, const char *action,
+							 RelFileLocator rlocator, BlockNumber blkno)
+{
+	/*
+	 * Only report once recovery has reached consistency. During crash
+	 * recovery, and before the consistency point of archive recovery, the
+	 * heap and VM pages can sit at different points of the WAL stream, so a
+	 * transient mismatch is expected and finishing replay resolves it.
+	 */
+	if (reachedConsistency && !heap_all_visible &&
+		(vm_oldbits & VISIBILITYMAP_ALL_VISIBLE))
+		ereport(WARNING,
+				(errcode(ERRCODE_DATA_CORRUPTED),
+				 errmsg("page %u of relation %s is not marked all-visible but its visibility map bit is set",
+						blkno, relpathperm(rlocator, MAIN_FORKNUM).str),
+				 errdetail("Redo is %s; the visibility map status was 0x%02X at record LSN %X/%X.",
+						   action, vm_oldbits, LSN_FORMAT_ARGS(lsn))));
+}
+
 /*
  * Clear visibility map bits for a single heap block during heap redo.
  *
@@ -35,8 +71,13 @@
  * 'heap_blkno' is the heap block whose VM bits should be cleared
  * 'wal_vm_block_id' is the WAL block reference id of the VM page
  * 'flags' specifies which visibility map bits to clear
+ *
+ * Returns the VM bits that were set for the block before this record cleared
+ * them, or 0 if we did not read them (no VM block reference, restored from a
+ * full-page image, or already up to date). Callers treat both the same: there
+ * is no evidence of divergence to report.
  */
-static void
+static uint8
 heap_xlog_vm_clear(XLogReaderState *record,
 				   RelFileLocator target_locator,
 				   BlockNumber heap_blkno,
@@ -44,9 +85,10 @@ heap_xlog_vm_clear(XLogReaderState *record,
 {
 	XLogRecPtr	lsn = record->EndRecPtr;
 	Buffer		vmbuffer = InvalidBuffer;
+	uint8		oldbits = 0;
 
 	if (!XLogRecHasBlockRef(record, wal_vm_block_id))
-		return;
+		return 0;
 
 	/*
 	 * If the vmbuffer was registered, use the recovery-specific routines to
@@ -63,12 +105,14 @@ heap_xlog_vm_clear(XLogReaderState *record,
 		if (PageIsNew(vmpage))
 			PageInit(vmpage, BLCKSZ, 0);
 
-		if (visibilitymap_clear(target_locator, heap_blkno, vmbuffer,
-								flags) & flags)
+		oldbits = visibilitymap_clear(target_locator, heap_blkno, vmbuffer, flags);
+		if (oldbits & flags)
 			PageSetLSN(vmpage, lsn);
 	}
 	if (BufferIsValid(vmbuffer))
 		UnlockReleaseBuffer(vmbuffer);
+
+	return oldbits;
 }
 
 /*
@@ -87,6 +131,13 @@ heap_xlog_prune_freeze(XLogReaderState *record)
 	uint8		vmflags = 0;
 	Size		freespace = 0;
 	bool		do_update_fsm = false;
+	XLogRedoAction heap_action;
+
+	/*
+	 * Whether the heap page was already marked all-visible before this
+	 * record. Only valid if heap_action is BLK_NEEDS_REDO.
+	 */
+	bool		heap_was_all_visible = false;
 
 	XLogRecGetBlockTag(record, 0, &rlocator, NULL, &blkno);
 	memcpy(&xlrec, maindataptr, SizeOfHeapPrune);
@@ -136,9 +187,10 @@ heap_xlog_prune_freeze(XLogReaderState *record)
 	 * If we have a full-page image of the heap block, restore it and we're
 	 * done with the heap block.
 	 */
-	if (XLogReadBufferForRedoExtended(record, 0, RBM_NORMAL,
-									  (xlrec.flags & XLHP_CLEANUP_LOCK) != 0,
-									  &buffer) == BLK_NEEDS_REDO)
+	heap_action = XLogReadBufferForRedoExtended(record, 0, RBM_NORMAL,
+												(xlrec.flags & XLHP_CLEANUP_LOCK) != 0,
+												&buffer);
+	if (heap_action == BLK_NEEDS_REDO)
 	{
 		Page		page = BufferGetPage(buffer);
 		OffsetNumber *redirected;
@@ -206,6 +258,8 @@ heap_xlog_prune_freeze(XLogReaderState *record)
 		/* There should be no more data */
 		Assert((char *) frz_offsets == dataptr + datalen);
 
+		heap_was_all_visible = PageIsAllVisible(page);
+
 		/*
 		 * The critical integrity requirement here is that we must never end
 		 * up with the visibility map bit set and the page-level
@@ -286,13 +340,31 @@ heap_xlog_prune_freeze(XLogReaderState *record)
 									  &vmbuffer) == BLK_NEEDS_REDO)
 	{
 		Page		vmpage = BufferGetPage(vmbuffer);
+		uint8		old_vmbits;
 
 		/* initialize the page if it was read as zeros */
 		if (PageIsNew(vmpage))
 			PageInit(vmpage, BLCKSZ, 0);
 
-		if (visibilitymap_set(blkno, vmbuffer, vmflags, rlocator) != vmflags)
+		old_vmbits = visibilitymap_set(blkno, vmbuffer, vmflags, rlocator);
+		if (old_vmbits != vmflags)
 			PageSetLSN(vmpage, lsn);
+
+		/*
+		 * If the all-visible bit was already set while the heap page was not
+		 * marked all-visible, the VM was corrupt on this standby. Redoing
+		 * this record has repaired it by marking the page all-visible; report
+		 * the pre-existing corruption.
+		 *
+		 * The VM block is read independently of the heap block and may need
+		 * redo even when the heap block was restored from a full-page image
+		 * or skipped by the LSN interlock. We only know the heap page's prior
+		 * state if we redid it.
+		 */
+		if (heap_action == BLK_NEEDS_REDO)
+			heap_xlog_warn_vm_corruption(heap_was_all_visible, old_vmbits, lsn,
+										 "marking the page all-visible",
+										 rlocator, blkno);
 	}
 
 	if (BufferIsValid(vmbuffer))
@@ -344,6 +416,7 @@ heap_xlog_delete(XLogReaderState *record)
 	BlockNumber blkno;
 	RelFileLocator target_locator;
 	ItemPointerData target_tid;
+	uint8		vm_oldbits = 0;
 
 	XLogRecGetBlockTag(record, HEAP_DELETE_BLKREF_HEAP, &target_locator, NULL,
 					   &blkno);
@@ -355,9 +428,9 @@ heap_xlog_delete(XLogReaderState *record)
 	 * already up-to-date.
 	 */
 	if (xlrec->flags & XLH_DELETE_ALL_VISIBLE_CLEARED)
-		heap_xlog_vm_clear(record, target_locator,
-						   blkno, HEAP_DELETE_BLKREF_VM,
-						   VISIBILITYMAP_VALID_BITS);
+		vm_oldbits = heap_xlog_vm_clear(record, target_locator,
+										blkno, HEAP_DELETE_BLKREF_VM,
+										VISIBILITYMAP_VALID_BITS);
 
 	if (XLogReadBufferForRedo(record, HEAP_DELETE_BLKREF_HEAP,
 							  &buffer) == BLK_NEEDS_REDO)
@@ -387,7 +460,12 @@ heap_xlog_delete(XLogReaderState *record)
 		PageSetPrunable(page, XLogRecGetXid(record));
 
 		if (xlrec->flags & XLH_DELETE_ALL_VISIBLE_CLEARED)
+		{
+			heap_xlog_warn_vm_corruption(PageIsAllVisible(page), vm_oldbits, lsn,
+										 "clearing the visibility map bits",
+										 target_locator, blkno);
 			PageClearAllVisible(page);
+		}
 
 		/* Make sure t_ctid is set correctly */
 		if (xlrec->flags & XLH_DELETE_IS_PARTITION_MOVE)
@@ -424,6 +502,7 @@ heap_xlog_insert(XLogReaderState *record)
 	BlockNumber blkno;
 	ItemPointerData target_tid;
 	XLogRedoAction action;
+	uint8		vm_oldbits = 0;
 
 	XLogRecGetBlockTag(record, HEAP_INSERT_BLKREF_HEAP, &target_locator, NULL,
 					   &blkno);
@@ -438,9 +517,9 @@ heap_xlog_insert(XLogReaderState *record)
 	 * already up-to-date.
 	 */
 	if (xlrec->flags & XLH_INSERT_ALL_VISIBLE_CLEARED)
-		heap_xlog_vm_clear(record, target_locator,
-						   blkno, HEAP_INSERT_BLKREF_VM,
-						   VISIBILITYMAP_VALID_BITS);
+		vm_oldbits = heap_xlog_vm_clear(record, target_locator,
+										blkno, HEAP_INSERT_BLKREF_VM,
+										VISIBILITYMAP_VALID_BITS);
 
 	/*
 	 * If we inserted the first and only tuple on the page, re-initialize the
@@ -503,7 +582,17 @@ heap_xlog_insert(XLogReaderState *record)
 		PageSetLSN(page, lsn);
 
 		if (xlrec->flags & XLH_INSERT_ALL_VISIBLE_CLEARED)
+		{
+			/*
+			 * A re-initialized page has no observable prior PD_ALL_VISIBLE,
+			 * so only check when we redid an existing page.
+			 */
+			if (!(XLogRecGetInfo(record) & XLOG_HEAP_INIT_PAGE))
+				heap_xlog_warn_vm_corruption(PageIsAllVisible(page), vm_oldbits, lsn,
+											 "clearing the visibility map bits",
+											 target_locator, blkno);
 			PageClearAllVisible(page);
+		}
 
 		MarkBufferDirty(buffer);
 	}
@@ -547,6 +636,7 @@ heap_xlog_multi_insert(XLogReaderState *record)
 	bool		isinit = (XLogRecGetInfo(record) & XLOG_HEAP_INIT_PAGE) != 0;
 	XLogRedoAction action;
 	Buffer		vmbuffer = InvalidBuffer;
+	uint8		vm_oldbits = 0;
 
 	/*
 	 * Insertion doesn't overwrite MVCC data, so no conflict processing is
@@ -570,9 +660,9 @@ heap_xlog_multi_insert(XLogReaderState *record)
 	 * all-visible in the VM while its PD_ALL_VISIBLE is clear.
 	 */
 	if (xlrec->flags & XLH_INSERT_ALL_VISIBLE_CLEARED)
-		heap_xlog_vm_clear(record, rlocator,
-						   blkno, HEAP_MULTI_INSERT_BLKREF_VM,
-						   VISIBILITYMAP_VALID_BITS);
+		vm_oldbits = heap_xlog_vm_clear(record, rlocator,
+										blkno, HEAP_MULTI_INSERT_BLKREF_VM,
+										VISIBILITYMAP_VALID_BITS);
 
 	if (isinit)
 	{
@@ -651,7 +741,13 @@ heap_xlog_multi_insert(XLogReaderState *record)
 		PageSetLSN(page, lsn);
 
 		if (xlrec->flags & XLH_INSERT_ALL_VISIBLE_CLEARED)
+		{
+			if (!isinit)
+				heap_xlog_warn_vm_corruption(PageIsAllVisible(page), vm_oldbits, lsn,
+											 "clearing the visibility map bits",
+											 rlocator, blkno);
 			PageClearAllVisible(page);
+		}
 
 		/*
 		 * XLH_INSERT_ALL_FROZEN_SET implies that all tuples are visible, so
@@ -772,6 +868,8 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 				npage;
 	bool		has_vm_old,
 				has_vm_new;
+	uint8		vm_old_oldbits = 0;
+	uint8		vm_new_oldbits = 0;
 	OffsetNumber offnum;
 	ItemId		lp;
 	HeapTupleData oldtup;
@@ -847,18 +945,21 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 			if (xlrec->flags & XLH_UPDATE_OLD_ALL_VISIBLE_CLEARED &&
 				visibilitymap_pin_ok(oldblk, vmbuffer_new))
 			{
-				if (visibilitymap_clear(rlocator, oldblk, vmbuffer_new,
-										VISIBILITYMAP_VALID_BITS))
+				vm_old_oldbits = visibilitymap_clear(rlocator, oldblk, vmbuffer_new,
+													 VISIBILITYMAP_VALID_BITS);
+				if (vm_old_oldbits)
 					PageSetLSN(BufferGetPage(vmbuffer_new), lsn);
 			}
 			/* If VM_NEW is registered, we are sure newblk is on VM_NEW */
-			if (visibilitymap_clear(rlocator, newblk, vmbuffer_new,
-									VISIBILITYMAP_VALID_BITS))
+			vm_new_oldbits = visibilitymap_clear(rlocator, newblk, vmbuffer_new,
+												 VISIBILITYMAP_VALID_BITS);
+			if (vm_new_oldbits)
 				PageSetLSN(BufferGetPage(vmbuffer_new), lsn);
 		}
 		if (BufferIsValid(vmbuffer_new))
 			UnlockReleaseBuffer(vmbuffer_new);
 	}
+
 	if (has_vm_old)
 	{
 		Buffer		vmbuffer_old = InvalidBuffer;
@@ -875,8 +976,9 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 			if (PageIsNew(vmpage))
 				PageInit(vmpage, BLCKSZ, 0);
 
-			if (visibilitymap_clear(rlocator, oldblk, vmbuffer_old,
-									VISIBILITYMAP_VALID_BITS))
+			vm_old_oldbits = visibilitymap_clear(rlocator, oldblk, vmbuffer_old,
+												 VISIBILITYMAP_VALID_BITS);
+			if (vm_old_oldbits)
 				PageSetLSN(BufferGetPage(vmbuffer_old), lsn);
 		}
 		if (BufferIsValid(vmbuffer_old))
@@ -931,7 +1033,12 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 		PageSetPrunable(opage, XLogRecGetXid(record));
 
 		if (xlrec->flags & XLH_UPDATE_OLD_ALL_VISIBLE_CLEARED)
+		{
+			heap_xlog_warn_vm_corruption(PageIsAllVisible(opage), vm_old_oldbits, lsn,
+										 "clearing the visibility map bits",
+										 rlocator, oldblk);
 			PageClearAllVisible(opage);
+		}
 
 		PageSetLSN(opage, lsn);
 		MarkBufferDirty(obuffer);
@@ -1053,7 +1160,19 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 			elog(PANIC, "failed to add tuple");
 
 		if (xlrec->flags & XLH_UPDATE_NEW_ALL_VISIBLE_CLEARED)
+		{
+			/*
+			 * Skip the check when the new tuple went onto the old page (the
+			 * old-page check above covers it) or onto a re-initialized page
+			 * (no observable prior PD_ALL_VISIBLE).
+			 */
+			if (oldblk != newblk &&
+				!(XLogRecGetInfo(record) & XLOG_HEAP_INIT_PAGE))
+				heap_xlog_warn_vm_corruption(PageIsAllVisible(npage), vm_new_oldbits, lsn,
+											 "clearing the visibility map bits",
+											 rlocator, newblk);
 			PageClearAllVisible(npage);
+		}
 
 		/* needed to update FSM below */
 		freespace = PageGetHeapFreeSpace(npage);
-- 
2.43.0

Reply via email to