On Sat, Sep 26, 2026 at 10:55 AM Xuneng Zhou <[email protected]> wrote: > > On Thu, Sep 24, 2026 at 7:34 PM Xuneng Zhou <[email protected]> wrote: > > > > On Thu, Sep 24, 2026 at 7:00 PM Xuneng Zhou <[email protected]> wrote: > > > > > > On Thu, Sep 24, 2026 at 5:28 PM Bertrand Drouvot > > > <[email protected]> wrote: > > > > > > > > Hi, > > > > > > > > On Thu, Sep 24, 2026 at 03:49:35PM +0800, Xuneng Zhou wrote: > > > > > Hi hackers, > > > > > > > > > > I don't see a clear solution to this potential issue, because the > > > > > interface > > > > > is a function, which means that the held snapshots cannot be popped > > > > > cleanly > > > > > since they belong to the surrounding executor. > > > > > > > > Thanks for the report and reproducers! > > > > > > > > Thinking out loud, I wonder if we could add a transient PGPROC state > > > > while a > > > > backend depends on recovery replay. After deadlock_timeout, > > > > ResolveRecoveryConflictWithVirtualXIDs() > > > > could check whether a VXID in its waitlist has that state set and, if > > > > so, use the > > > > existing recovery conflict cancellation path. > > > > > > > > Does that make sense to you and others? If so, I can have a look at > > > > preparing a > > > > patch. > > > > > > The overall direction looks promising to me, and I haven't come up > > > with a simpler fix. As a side benefit, it could also break the > > > potential deadlock where 'WAIT' command waits for replay while > > > recovery waits for the same backend's VXID. It would be good to hear > > > more echo before heading to implementation. > > > > Here's the reproducer for the mentioned VXID issue. I think we need to > > test the fix for it as well, since the underlying issue remains the > > same. The reproducers could fit in existing test files like 031, but > > for clarity, they are in standalone files. Also CCed Alexander for > > this. > > After more investigation, both slot functions seem also vulnerable to > VXID deadlock issue like the WAIT command. My original thought for the > fix of the issue is to let ResolveRecoveryConflictWithSnapshot make > the blocking decision based on the actual snapshot conflict rather > than simply checking whether the VXID has gone away. However, that is > more complex and needs more consideration than what you proposed. It > could be a follow-up optimization, not necessarily the bug fix. > Another problem that both functions suffered is the heavyweight > deadlock issue[1]. > > I also asked Astra to do a broader inspection for the same categorical > issue in the tree, and it did find more, which I'll share later. > > [1] > https://www.postgresql.org/message-id/CABPTF7U0gW5%2B-4oL7-qdML-yerZxUb7ku4QXp7JxCYo0qyJ_Tw%40mail.gmail.com
Hold on a bit. The deeper I dig, the more interesting it gets. I'll share something very different soon.. -- Regards, Xuneng Zhou HighGo Software Co., Ltd.
