wudidapaopao opened a new pull request, #25990:
URL: https://github.com/apache/datafusion/pull/25990

   ## Which issue does this PR close?
   
   - Closes #25989.
   
   ## Rationale for this change
   
   `COUNT(non_nullable_column)` counts every input row, but evaluating it still 
reads and materializes the column unnecessarily.
   
   ## What changes are included in this PR?
   
   - Rewrite safe, non-`DISTINCT` counts over non-nullable columns to 
`COUNT(1)` during expression simplification while preserving output names and 
schemas.
   - Keep `SingleDistinctToGroupBy` rewrites working when name preservation 
adds an alias around the rewritten count.
   - Avoid marking unchanged aggregate simplifications as transformed.
   
   ## What is the testing strategy for this PR?
   
   Covered by unit tests for nullable, non-nullable, multi-argument, 
expression, and `DISTINCT` counts, plus optimizer, DataFrame, object-store 
column pruning, unparser, and sqllogictest coverage.
   
   Validated with:
   
   - `cargo fmt --all -- --check`
   - `cargo clippy --all-targets --all-features -- -D warnings`
   - aggregate and optimizer unit/integration tests
   - targeted core integration tests
   - the full default and TPC-H sqllogictest suites
   
   ## Are there any user-facing changes?
   
   Query results and output schemas are unchanged. Optimized plans can show 
`count(1) AS count(column)`, and scans can prune non-nullable columns used only 
by `COUNT`.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to