[
https://issues.apache.org/jira/browse/SPARK-59426?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-59426:
-----------------------------------
Labels: pull-request-available (was: )
> Avoid per-lookup allocation in EquivalentExpressions.ExpressionEquals.hashCode
> ------------------------------------------------------------------------------
>
> Key: SPARK-59426
> URL: https://issues.apache.org/jira/browse/SPARK-59426
> Project: Spark
> Issue Type: Improvement
> Components: SQL
> Affects Versions: 4.1.0
> Reporter: David Mollitor
> Priority: Trivial
> Labels: pull-request-available
>
> h2. Summary
> {{ExpressionEquals}} is the key type of
> {{EquivalentExpressions.equivalenceMap}}
> ({{mutable.HashMap[ExpressionEquals, ExpressionStats]}}), the map that drives
> common-subexpression elimination (CSE). Its {{hashCode}} was:
> {code:scala}
> override def hashCode: Int = Objects.hash(e.semanticHash(): Integer, height:
> Integer)
> {code}
> {{java.util.Objects.hash(...)}} is varargs, so every call boxes both ints to
> {{Integer}} and
> allocates an {{Object[2]}}. This key is hashed on every {{addExpr}} /
> {{updateExprInMap}} /
> {{getExprState}} lookup during CSE, so expression-heavy queries generate a
> steady stream of short-lived allocations during optimization/codegen.
> While here, reorder {{equals}} to test the O(1) {{height}} before the
> recursive {{semanticEquals}}. {{semanticEquals}} implies equal {{height}}, so
> an unequal height short-circuits without walking the expression tree; {{&&}}
> yields the identical boolean. {{equals}} runs only on hash-bucket collisions,
> which is exactly where the cheap pre-check pays off.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]