sunchao commented on PR #6427:
URL: 
https://github.com/apache/datafusion-comet/pull/6427#issuecomment-5916719828

   @viirya You’re right about Spark’s automatically injected runtime filters. I 
traced the original case, and my explanation was incorrect: **the literal comes 
from application code, not from Spark materializing a scalar subquery**.
   
   The application builds and serializes a Bloom filter, constructs 
`BloomFilterMightContain(Literal(bytes, BinaryType), key)`, and passes that 
expression through `DataFrame.filter`. That literal survives into 
`CometFilterExec.condition`.
   
   I reproduced this through an actual DataFrame query with native Comet 
execution, rather than manually constructing a physical filter. A 
1,048,588-byte serialized Bloom produced a 2,097,226-character `SparkPlanInfo` 
string, while returning the expected results. The automatic runtime-filter 
controls retained ScalarSubquery with AQE both on and off.
   
   So the display issue is reachable, but the rationale should specifically 
describe application-built Bloom predicates. I’ll correct the description and 
add a real-query regression covering that path.
   
   I also agree that expression-level formatting in Spark would provide broader 
coverage, including fallback and projection cases. This patch is narrower: it 
summarizes these literals only in Comet filter displays.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to