Hi,

1. Update the pg_stats_ext_exprs view to expose STATISTIC_KIND_MCV_VALUE_SORTED
through most_common_vals and most_common_freqs.

2. Regarding pg_statistic_get_difference():

The input statistics can come either from pg_stats or directly from the user. 
I initially considered improving the detection of the stat kind during import, 
but there are several considerations and implementation difficulties. 
For now, I think it is better to keep the existing logic.

* Adding a stat kind parameter would rely on the caller to make the correct 
decision, 
     and would also affect quite a few interfaces.
* Another idea is to inspect most_common_vals and most_common_freqs during 
import, 
sort them if possible, and generate STATISTIC_KIND_MCV_VALUE_SORTED; otherwise, 
keep STATISTIC_KIND_MCV.

I looked into the latter, but it seems more complicated than expected because 
both inputs are text. 
We would need to deserialize them into the actual data type, sort them, and 
serialize them again. 
I have not worked out the details yet. If anyone is familiar with this part of 
the code and has suggestions, 
I would be happy to investigate further.

3. I think pg_statistic_get_difference() could also have stronger validation in 
the future.

For input from pg_stats, the core code has already processed the statistics, so 
for STATISTIC_KIND_MCV, 
most_common_freqs[0] is the maximum and most_common_freqs[n - 1] is the minimum.

For user-provided input, however, there is currently no such restriction. We 
only perform basic checks, 
such as array lengths and NULLs. We could consider strengthening these checks 
in the future.


regards,
--
ZizhuanLiu (X-MAN) 
[email protected]

Attachment: v7-0001-Optimize-MCV-statistics-for-sortable-types.patch
Description: Binary data

Reply via email to