Messages by Thread
-
Re: [I] datafusion-proto: logical plan decode re-normalizes already-normalized plans, superlinear on wide plans [datafusion]
via GitHub
-
[I] `array_concat` on nested lists panics or fails with type mismatch errors [datafusion]
via GitHub
-
[I] `nvl` / `ifnull` convert decimals to Float64 and timestamps and dates to strings [datafusion]
via GitHub
-
[I] `decimal % 1` returns 0 for non-nullable decimal columns [datafusion]
via GitHub
-
Re: [PR] perf: Add an adaptive integer membership prefilter for hash joins [datafusion]
via GitHub
-
[PR] fix: defer Parquet conversion errors until a row group is decoded, as Spark does [datafusion-comet]
via GitHub
-
Re: [D] DISCUSSION: San Francisco DataFusion Meetup November 2026 [datafusion]
via GitHub
-
[I] SQL unparser drops the time zone from timestamp literals [datafusion]
via GitHub
-
Re: [I] array_distinct and array_union diverge from Spark on -0.0 for Spark versions without SPARK-54918 [datafusion-comet]
via GitHub
-
[PR] fix(proto): preserve EmptyRelation schema across logical plan round trip [datafusion]
via GitHub
-
[I] Queries with multiple `SUM(expr + literal)` aggregates fail or change result type [datafusion]
via GitHub
-
[I] Spark `negative` returns wrong results for decimal columns [datafusion]
via GitHub
-
[I] Comparing a timestamp cast to a time zone against a literal returns wrong results [datafusion]
via GitHub
-
[PR] feat: Add option to assign Parquet row groups to file ranges by midpoint [datafusion]
via GitHub
-
[PR] docs: add community showcase volumes 4 and 5 [datafusion]
via GitHub
-
[I] `convert_filters_to_predicate` negates a weakened AND, producing a predicate that drops matching rows [datafusion-iceberg]
via GitHub
-
[I] LIKE prefix pushdown ignores backslash escapes and drops matching rows [datafusion-iceberg]
via GitHub
-
Re: [I] Native Iceberg write merges -0.0 and 0.0 rows into one float/double identity partition [datafusion-comet]
via GitHub
-
[I] Option to assign Parquet row groups to file ranges by midpoint, as Spark does [datafusion]
via GitHub
-
[I] `SELECT *` fails after a column is added until the next write to the table [datafusion-iceberg]
via GitHub
-
[I] INSERT into an unpartitioned table fails when fanout is disabled [datafusion-iceberg]
via GitHub
-
[I] Partitioned INSERT computes partition values from the wrong columns when the input plan ends in a projection [datafusion-iceberg]
via GitHub
-
[I] Clustered (fanout disabled) writes fail on sorted input because batches are split in HashMap order [datafusion-iceberg]
via GitHub
-
[I] INSERT into a partitioned table with fanout disabled always fails: the optimizer removes the partition sort [datafusion-iceberg]
via GitHub
-
[I] Metadata tables ignore the scan projection, so projected and aggregate queries return wrong columns or fail [datafusion-iceberg]
via GitHub
-
[I] Filter pushdown strips casts on columns, which drops matching rows or fails valid queries [datafusion-iceberg]
via GitHub
-
[I] `CREATE TABLE` / `DROP TABLE` hang on a current-thread tokio runtime and panic outside a runtime [datafusion-iceberg]
via GitHub
-
[I] datafusion-proto: EmptyRelation loses its schema on a logical plan round trip [datafusion]
via GitHub
-
Re: [I] General framework to decorrelate the subqueries [datafusion]
via GitHub
-
[I] Grouped `count(DISTINCT x)` on numbers takes a slower path [datafusion]
via GitHub
-
[I] Support native existence sort-merge joins (blocked on DataFusion mark-SMJ output buffering) [datafusion-comet]
via GitHub
-
[PR] fix: prefer Unsupported when combining nested cast support [datafusion-comet]
via GitHub
-
Re: [I] Correlated `EXISTS` subquery with `OFFSET` returns wrong results [datafusion]
via GitHub
-
[PR] docs: check for behavior changes against the latest release in the review skill [datafusion-comet]
via GitHub
-
[PR] chore: Upgrade to Rust 1.99, MSRV to 1.95 [datafusion]
via GitHub
-
[I] Native Parquet scan assigns row groups to splits differently from Spark [datafusion-comet]
via GitHub
-
[PR] chore(deps): bump pyjwt from 2.13.0 to 2.15.0 in /python [datafusion-ballista]
via GitHub
-
[I] Record whether a window frame was set explicitly, so `ExprFunctionExt` can build on an existing window function [datafusion]
via GitHub
-
Re: [PR] fix: keep the recovery projection when leaf pushdown would drop a computed same-name column [datafusion]
via GitHub
-
Re: [I] Wrong results: leaf expression pushdown removes a computed column that has the same name as its input column [datafusion]
via GitHub
-
Re: [PR] feat: add parquet_file_metadata and parquet_page_index to datafusion-cli [datafusion]
via GitHub
-
Re: [PR] refactor: Use cast preimages for cast predicate rewrites [datafusion]
via GitHub
-
Re: [I] Specialized GroupValues Implementation for Ordered Data [datafusion]
via GitHub
-
Re: [I] Comet 1.1.0 Release (September) [datafusion-comet]
via GitHub
-
Re: [PR] fix: support zero-field struct group keys in aggregate emit [datafusion]
via GitHub
-
[I] `sum(DISTINCT x)` with `GROUP BY` uses twice the memory or more, unless the query also has `count(*)` [datafusion]
via GitHub
-
[PR] fix: [branch-1.1] fall back to Spark for _metadata.file_block_start and file_block_length (#6510) [datafusion-comet]
via GitHub
-
[PR] fix: fall back to Spark for _metadata.file_block_start and file_block_length [datafusion-comet]
via GitHub
-
[PR] docs: add AGENTS.md for coding agents [datafusion-iceberg]
via GitHub
-
[I] why python version older than the rust version [datafusion-python]
via GitHub
-
[PR] feat: reuse S3 credentials until shortly before their reported expiry [datafusion-comet]
via GitHub
-
[I] Quoted AST values and tokens can serialize to non-equivalent SQL [datafusion-sqlparser-rs]
via GitHub
-
[I] Add an AGENTS.md like other DataFusion repositories [datafusion-iceberg]
via GitHub
-
[I] S3 credential SPI: honor expirationEpochMillis consistently on the Parquet and Iceberg paths [datafusion-comet]
via GitHub
-
Re: [PR] fix: reject scalar subqueries inside a codegen-dispatch kernel [datafusion-comet]
via GitHub
-
[PR] chore(deps): bump urllib3 from 2.7.0 to 2.8.0 in /python [datafusion-ballista]
via GitHub
-
Re: [PR] feat: route map lookups through codegen dispatcher [datafusion-comet]
via GitHub
-
[I] Crate name `datafusion-iceberg` is taken on crates.io; switch back to `iceberg-datafusion` [datafusion-iceberg]
via GitHub
-
[I] Add release infrastructure [datafusion-iceberg]
via GitHub
-
[PR] fix: build with Rust 1.99, which deprecates the legacy f64 constants and fetch_update [datafusion-comet]
via GitHub
-
[I] Native Parquet scan returns wrong _metadata.file_block_start when Spark splits a file [datafusion-comet]
via GitHub
-
Re: [PR] perf: optimize JVM columnar-to-row conversion [datafusion-comet]
via GitHub
-
[I] Native Parquet scan fails on an empty file whose nested field type differs from the read schema [datafusion-comet]
via GitHub
-
[PR] chore: add release infrastructure and rename crate back to `iceberg-datafusion` [datafusion-iceberg]
via GitHub
-
Re: [PR] feat: custom Rust UDFs via arrow-ffi [experimental] [datafusion-comet]
via GitHub
-
Re: [I] Examples of using `TreeNode` APIs to walk and manipulate LogicalPlans [datafusion]
via GitHub
-
Re: [PR] feat: cost-based join order enumeration [datafusion]
via GitHub
-
[I] Native Iceberg scan fails on files written before a nested field was added, and since 1.1.0 null checks and explode hit it too [datafusion-comet]
via GitHub
-
[PR] refactor: move the self-contained operators from core into a new operators crate [datafusion-comet]
via GitHub
-
Re: [PR] test: close five mutation-testing gaps in leaf expression pushdown guards [datafusion]
via GitHub
-
Re: [I] Physical optimizer re-runs rules on plans they have already settled [datafusion]
via GitHub
-
[PR] fix: read the Iceberg write gate's storage scheme the same way the native factory does [datafusion-comet]
via GitHub
-
Re: [PR] fix: show results in the nine silent top-level examples [datafusion-python]
via GitHub
-
Re: [I] Split physical optimizer rules into enforcement and optimization phases, with a convergence loop for the optimizers [datafusion]
via GitHub
-
Re: [PR] perf(aggregate): specialize fully ordered group keys without hashing [datafusion]
via GitHub
-
[PR] fix: charge semi/anti sort-merge join key groups by their sliced size [datafusion]
via GitHub
-
[PR] Bump bigdecimal from 0.4.10 to 0.4.11 [datafusion-sqlparser-rs]
via GitHub
-
Re: [I] Make `AggregateExec` state modeling and public updates safe [datafusion]
via GitHub
-
[I] Experimental local execution mode that runs a whole query as one DataFusion graph [datafusion-comet]
via GitHub
-
[PR] perf: add fully ordered group values implementation for primitive [datafusion]
via GitHub
-
[PR] docs: show array_distinct keeping first-seen order [datafusion]
via GitHub
-
Re: [PR] fix: keep the OFFSET of a correlated EXISTS subquery [datafusion]
via GitHub
-
[PR] y [datafusion]
via GitHub
-
Re: [I] `CometCast.isSupported` returns the first non-Compatible child, so `Incompatible` can mask `Unsupported` [datafusion-comet]
via GitHub
-
[PR] feat: Experimental local execution mode that runs a whole query as one DataFusion graph [datafusion-comet]
via GitHub
-
[I] Implement optimizer hook for combining partial/final aggregate in `AggregateExec` [datafusion]
via GitHub
-
[PR] fix: compact oversized nested list_extract buffers [datafusion-comet]
via GitHub
-
[PR] test: cover typed Dataset filters over sliced aggregate batches [datafusion-comet]
via GitHub
-
Re: [I] Push Down Offset to TableScan [datafusion]
via GitHub
-
[PR] chore(deps): bump pyjwt from 2.12.0 to 2.15.0 [datafusion-sandbox]
via GitHub
-
[PR] docs: add repository README [datafusion-iceberg]
via GitHub
-
Re: [PR] feat: admit string maps in Spark-to-Comet conversion [datafusion-comet]
via GitHub
-
Re: [I] Native implementation of `get_json_object` returns last value for duplicate keys, Spark returns first [datafusion-comet]
via GitHub
-
Re: [PR] perf(proto): avoid re-normalizing logical plans [datafusion]
via GitHub
-
[PR] fix: redact credentials in object-store URL diagnostics [datafusion]
via GitHub
-
[PR] perf: reduce Parquet runtime filter schema guard overhead [datafusion-comet]
via GitHub
-
Re: [PR] perf: avoid quadratic planning for SELECTs with many aggregates [datafusion]
via GitHub
-
[PR] build(deps): bump pyjwt from 2.13.0 to 2.15.0 [datafusion-python]
via GitHub
-
[PR] test: share collation coverage across Spark 4.x [datafusion-comet]
via GitHub
-
[PR] chore: move native write configs under spark.comet.write [datafusion-comet]
via GitHub
-
[I] Avoid redundant sorts for scalar subquery expressions [datafusion]
via GitHub
-
[I] Add checked time arithmetic regression coverage for scalar UDF type recovery [datafusion]
via GitHub
-
Re: [PR] test: enable SPARK-57298 collect_set tests [datafusion-comet]
via GitHub
-
Re: [I] native: `JVMClasses::with_env: JAVA_VM not initialized` aborts a multi-suite JVM [datafusion-comet]
via GitHub
-
[PR] fix: load the bundled native library only once per class loader [datafusion-comet]
via GitHub
-
Re: [PR] chore(deps): bump setuptools from 82.0.0 to 83.0.0 [datafusion-sandbox]
via GitHub
-
Re: [PR] chore(deps-dev): bump shell-quote from 1.8.3 to 1.10.0 in /datafusion/wasmtest/datafusion-wasm-app [datafusion-sandbox]
via GitHub
-
Re: [PR] feat: Refactor NLJ into an extensible framework for specialized joins [datafusion]
via GitHub
-
Re: [PR] fix: use build_join_schema for CrossJoinExec output schema metadata [datafusion]
via GitHub
-
[PR] test: [branch-1.1] cover sliced boolean arrays in explode (#6473) [datafusion-comet]
via GitHub
-
[PR] fix: [branch-1.1] fall back for incompatible regression aggregates (#6451) [datafusion-comet]
via GitHub
-
[PR] fix: [branch-1.1] rescale decimals in generated dispatchers (#6455) [datafusion-comet]
via GitHub
-
[PR] fix: [branch-1.1] coerce native IF branches to a common type (#6458) [datafusion-comet]
via GitHub
-
[PR] fix: [branch-1.1] make adaptive aggregation skipping opt-in (#6474) [datafusion-comet]
via GitHub
-
[PR] fix: [branch-1.1] fall back for rank limits over nested float keys (#6468) [datafusion-comet]
via GitHub
-
[PR] fix: [branch-1.1] match pre-epoch Iceberg temporal rounding (#6456) [datafusion-comet]
via GitHub
-
[PR] fix: Fix make_array null input handling [datafusion]
via GitHub
-
Re: [I] [Doc] CAST collated-string handling on Spark 4.0+ is implicit and untested [datafusion-comet]
via GitHub
-
[I] Assess native (Rust) code generation for fused expression evaluation [datafusion-comet]
via GitHub
-
Re: [I] Nested duplicate names in a metadata-free Parquet file bypass the native resolver when the file schema equals the requested schema [datafusion-comet]
via GitHub
-
[I] Filtered semi/anti sort-merge join is slow on small key groups and charges each group a whole batch [datafusion]
via GitHub
-
Re: [PR] fix: match Spark statistical aggregate updates and merges [datafusion-comet]
via GitHub