[
https://issues.apache.org/jira/browse/SPARK-59056?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-59056:
-----------------------------------
Labels: pull-request-available (was: )
> Pre-size HashSets in ParquetFilters IN/InSet predicate pushdown
> ---------------------------------------------------------------
>
> Key: SPARK-59056
> URL: https://issues.apache.org/jira/browse/SPARK-59056
> Project: Spark
> Issue Type: Improvement
> Components: SQL
> Affects Versions: 4.1.0
> Reporter: David Mollitor
> Priority: Minor
> Labels: pull-request-available
>
> h2. Summary
> {{ParquetFilters.makeInPredicate}} builds a {{java.util.HashSet}} for every
> {{IN}} predicate pushed down to Parquet, without an initial capacity. The
> final element count is exactly {{values.length}} and is known before the set
> is populated, so the set can be pre-sized to avoid internal rehashing and the
> intermediate bucket-array allocations.
> h2. Details
> {{makeInPredicate}} has 12 near-identical sites of the form:
> {code:scala}
> val set = new HashSet[T]()
> for (value <- values) {
> set.add(...) // one element per entry in `values`
> }
> FilterApi.in(column(n), set)
> {code}
> Each {{HashSet}} starts at the default capacity (16) and rehashes as it grows
> toward
> {{values.length}} elements. This set is only built when the {{IN}} list size
> exceeds
> {{spark.sql.parquet.pushdown.inFilterThreshold}} (default 10); smaller lists
> take a per-element equality OR-chain and never build a {{{}HashSet{}}}. So by
> construction the set is always populated with more than the threshold's worth
> of values, and for larger lists (dozens to thousands – generated {{IN}}
> lists, dynamic filters) the default-capacity set rehashes several times per
> pushed-down predicate, on every Parquet file split scanned.
> h2. Proposed change
> Pre-size each set from the known element count using Guava's
> {{Sets.newHashSetWithExpectedSize(values.length)}} (already a Spark
> dependency; consistent with the existing {{Maps.newHashMapWithExpectedSize}}
> usage in {{{}catalyst{}}}). The Guava helper is used rather than {{new
> HashSet(values.length)}} because {{{}java.util.HashSet{}}}'s int constructor
> treats the argument as bucket capacity, not expected size, so the naive form
> would still rehash at ~75% load.
> The change is capacity-only and behavior-preserving – the resulting predicate
> semantics are identical.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]