On 25/09/2026 4:21 AM, Jeff Davis wrote:
On Thu, 2026-09-24 at 21:30 +0300, Konstantin Knizhnik wrote:
v2 keeps flushedUpto monotonic.  Startup uses a separate
applyFlushedUpto for "is this streamed WAL readable?"
The reread-from-primary behavior seems to trace to 6da07cd80d, which
looks more like an opportunistic retry than a policy.

We could make it a policy to try to get readable WAL from the primary,
and then keep the extra state you propose. But I think we need more
explanation about why we'd expect this situation (corrupt and flushed
on standby and valid on primary) to occur.

For instance:

  /* ...Only after a failed read of
   * WAL that flushedUpto still reports as present do we
   * rewind the apply pointer, so startup waits for
   * replacement bytes.

raises questions for anyone reading that code about why it might be
flushed and unreadable.

Regards,
        Jeff Davis


Paul has found and fixed original problem at primary side:

https://www.postgresql.org/message-id/178956158235.97809.18200289969141907831%40mail.gmail.com

But the proposed patch is still worth keeping as a standby self-heal for any other invalid bytes under a high |flushedUpto| (other producer bugs, CF 5199-style missing file, on-disk junk) still hang without it. The included test demonstrates it.

Reply via email to