On 25/09/2026 4:21 AM, Jeff Davis wrote:
On Thu, 2026-09-24 at 21:30 +0300, Konstantin Knizhnik wrote:
v2 keeps flushedUpto monotonic. Startup uses a separate
applyFlushedUpto for "is this streamed WAL readable?"
The reread-from-primary behavior seems to trace to 6da07cd80d, which
looks more like an opportunistic retry than a policy.
We could make it a policy to try to get readable WAL from the primary,
and then keep the extra state you propose. But I think we need more
explanation about why we'd expect this situation (corrupt and flushed
on standby and valid on primary) to occur.
For instance:
/* ...Only after a failed read of
* WAL that flushedUpto still reports as present do we
* rewind the apply pointer, so startup waits for
* replacement bytes.
raises questions for anyone reading that code about why it might be
flushed and unreadable.
Regards,
Jeff Davis
Paul has found and fixed original problem at primary side:
https://www.postgresql.org/message-id/178956158235.97809.18200289969141907831%40mail.gmail.com
But the proposed patch is still worth keeping as a standby self-heal for
any other invalid bytes under a high |flushedUpto| (other producer bugs,
CF 5199-style missing file, on-disk junk) still hang without it. The
included test demonstrates it.