On Tue, Aug 18, 2026 at 6:22 PM Ayush Tiwari
<[email protected]> wrote:
> I tried a case with 1000 unmatched hash tuples where the filter allowed
> only one through. The executor examined all 1000 and reported "Rows
> Removed by Filter: 999", while path->rows was one. The patch reduced the
> cost by 9.99, exactly 999 * cpu_tuple_cost.
>
> Could using path->rows therefore undercharge the rows that were examined
> but filtered out?

Yeah, you're right.  path->rows is the row count after all the quals
have been applied, while the per-tuple charges here are meant to cover
the rows that come out of the hash-clause matching, before any other
quals filter them.

I think a more accurate way is to compute hashjointuples from the
inner side directly:

  right-semi: inner_path_rows * outer_match_frac
  right-anti: inner_path_rows * (1 - outer_match_frac)

This is mirroring what JOIN_SEMI and JOIN_ANTI already do with the
outer side.  Both cpu_tuple_cost and the quals not used for hashing
are charged on that count, so the 1000 rows your executor examined are
all charged, and the filter's selectivity affects only the join's
output row estimate.

Attached v3 patch does that.

- Richard

Attachment: v3-0001-Fix-CPU-cost-of-right-semi-and-right-anti-hash-jo.patch
Description: Binary data

Reply via email to