Messages by Thread
-
Re: [PR] fix(scheduler): retire AQE stages cancelled by replan and discard late tasks [datafusion-ballista]
via GitHub
-
[PR] fix(release): prepare the 55 release process [datafusion-ballista]
via GitHub
-
[I] Native Iceberg writer counts NaNs under NULL structs and NULL list or map entries [datafusion-comet]
via GitHub
-
[PR] fix(docker): publish images from release tag pushes [datafusion-ballista]
via GitHub
-
Re: [I] Accelerated mapInArrow reads Python output at the declared types without checking them [datafusion-comet]
via GitHub
-
Re: [PR] fix: keep the running hash for nested dictionaries in Spark xxhash64 [datafusion]
via GitHub
-
[PR] fix: [branch-1.1] gate array distinct and union signed-zero semantics by Spark version (#5750) [datafusion-comet]
via GitHub
-
Re: [I] Support more types for `approx_distinct` function [datafusion]
via GitHub
-
Re: [PR] feat: Arrow Flight SQL frontend for the scheduler + ADBC support [datafusion-ballista]
via GitHub
-
[PR] fix: [branch-1.1] fall back from the native Iceberg scan when a nested field was added or renamed (#6543) [datafusion-comet]
via GitHub
-
[I] Tighten integer division interval bounds when the divisor has a zero endpoint [datafusion]
via GitHub
-
[PR] fix(functions): log does not preserve the order of inputs that may be negative [datafusion]
via GitHub
-
Re: [I] Potential Optimizations for `hash_union_array` [datafusion]
via GitHub
-
[I] Parallel Parquet writes delay reporting column-worker failures [datafusion]
via GitHub
-
[PR] Add regression tests for nullable NOT IN expressions [datafusion]
via GitHub
-
[PR] doc: add comments for `LocalLimitExec` and `GlobalLimitExec` [datafusion]
via GitHub
-
Re: [I] Improve shuffle (column) statistics [datafusion-ballista]
via GitHub
-
Re: [PR] IN LIST: retain short lists with specialized filters [datafusion]
via GitHub
-
Re: [I] datafusion-proto: column qualifiers containing `.` are silently corrupted on round-trip [datafusion]
via GitHub
-
[PR] feat: rewrite count over non-nullable columns [datafusion]
via GitHub
-
[I] Rewrite COUNT over non-nullable columns to COUNT(1) [datafusion]
via GitHub
-
[PR] fix: keep grouped integer DISTINCT counts on the native accumulator path [datafusion]
via GitHub
-
[PR] fix: convert native ReadAncientDatetime errors to SparkUpgradeException [datafusion-comet]
via GitHub
-
Re: [D] How DataFusion could support other compute engines (libcudf, velox) [datafusion]
via GitHub
-
Re: [I] Support pruning on INT96 timestamp columns via `SortOrder::INT96_TIMESTAMP` / `ColumnOrder::INT96_TIMESTAMP_ORDER` [datafusion]
via GitHub
-
[PR] feat(parquet): prune explicitly ordered INT96 timestamps [datafusion]
via GitHub
-
[PR] fix: release semi/anti/mark sort-merge join reservations before the final output [datafusion]
via GitHub
-
[PR] fix: handle zero endpoints in division interval bounds [datafusion]
via GitHub
-
[I] Native rebase refusals raise CometNativeException instead of SparkUpgradeException [datafusion-comet]
via GitHub
-
[PR] perf: materialize completed groups once in OrderedSingleAggregateStream [datafusion]
via GitHub
-
Re: [PR] feat(sort-shuffle): accept Option<Partitioning> [datafusion-ballista]
via GitHub
-
[PR] fix: keep aggregate limits when combining or rebuilding AggregateExec [datafusion]
via GitHub
-
Re: [PR] feat: add docker-compose.quick.yml and fix onboarding docs [datafusion-ballista]
via GitHub
-
Re: [PR] feat: add %%sql --limit and --no-display support [datafusion-ballista]
via GitHub
-
Re: [PR] chore(deps-dev): bump webpack-dev-server from 5.2.1 to 5.2.6 in /datafusion/wasmtest/datafusion-wasm-app [datafusion-sandbox]
via GitHub
-
Re: [PR] fix: avoid dictionary key overflow in group value output [datafusion]
via GitHub
-
Re: [PR] Update contributor guidelines regarding AI spam, reviews, code [datafusion]
via GitHub
-
[I] testing [datafusion]
via GitHub
-
Re: [PR] fix(cli): respect `sql_parser.recursion_limit` when parsing and validating input [datafusion]
via GitHub
-
[PR] fix: run sync catalog calls on a separate runtime [datafusion-iceberg]
via GitHub
-
[PR] docs: note that the peak native aggregate memory metric is not reported [datafusion-comet]
via GitHub
-
Re: [PR] fix: reject invalid placeholders in CREATE FUNCTION bodies at definition time [datafusion]
via GitHub
-
[PR] fix: return -0.0 for signum(-0.0), as Spark does [datafusion-comet]
via GitHub
-
[I] spark: array_contains compares floats by bits, where Spark treats -0.0 as 0.0 and all NaNs as equal [datafusion]
via GitHub
-
Re: [I] Missing Redshift constructs [datafusion-sqlparser-rs]
via GitHub
-
Re: [I] Discussion: guidelines for LLM-generated PR reviews [datafusion]
via GitHub
-
[PR] [branch-55] fix: preserve compound expression equivalences through projection (#25922) [datafusion]
via GitHub
-
Re: [I] `AND` chain pre-selection tests its threshold against the accumulated prefix, not the next conjunct [datafusion]
via GitHub
-
[I] Support OneRowRelation as a native Comet source [datafusion-comet]
via GitHub
-
Re: [PR] feat: Use DataFusion date_trunc for scalar-format timestamp truncation [datafusion-comet]
via GitHub
-
[PR] feat: add an experimental Comet version of RangeExec [datafusion-comet]
via GitHub
-
[PR] fix: [branch-1.1] dispatch StaticInvoke and Invoke only into Spark's own classes (#6542) [datafusion-comet]
via GitHub
-
Re: [PR] Update contributor guidelines regarding AI spam [datafusion]
via GitHub
-
Re: [PR] The great `Box`-ification experiment [datafusion-sqlparser-rs]
via GitHub
-
Re: [PR] feat: support SortAggregateExec [datafusion-comet]
via GitHub
-
[PR] fix: stop struct outputs of the codegen dispatcher from leaking Arrow memory [datafusion-comet]
via GitHub
-
Re: [PR] feat: add build-gated Hudi read integration [datafusion-ballista]
via GitHub
-
Re: [PR] feat: add Hudi contrib build gate [datafusion-ballista]
via GitHub
-
[PR] docs: explain which Spark operators Comet leaves in place [datafusion-comet]
via GitHub
-
[PR] improve Spark from_utc_timestamp compatibility [datafusion]
via GitHub
-
[I] Comet 1.2.0 Release [datafusion-comet]
via GitHub
-
[I] Native map construction doesn't match Spark 4.0+ float key normalization (-0.0 keys, missing DUPLICATED_MAP_KEY) [datafusion-comet]
via GitHub
-
[PR] feat: support native wide-decimal hashing [datafusion-comet]
via GitHub
-
[PR] perf: read the shuffle natively in operators AQE reuses from the initial plan [datafusion-comet]
via GitHub
-
[I] Internal error: GROUP BY over UNION ALL fails when a branch projects `coalesce(<nullable bool>, FALSE)` (CSE rewrite is physically nullable) [datafusion]
via GitHub
-
[I] Native Iceberg scan returns NULL for a nested field renamed after its data file was written [datafusion-comet]
via GitHub
-
[PR] feat: [branch-1.1] reuse S3 credentials until shortly before their reported expiry (#6509) [datafusion-comet]
via GitHub
-
[PR] fix: let a spilled final aggregate read its spill files back past its memory share [datafusion-comet]
via GitHub
-
Re: [I] Decide keep-or-lift for each remaining native Iceberg write eligibility restriction [datafusion-comet]
via GitHub
-
Re: [I] Call for Presentations: Community Showcase: Regular series for sharing what you're building with DataFusion [datafusion]
via GitHub
-
Re: [PR] fix: align wide-decimal ANSI overflow value with Spark [datafusion-comet]
via GitHub
-
Re: [PR] test: enable `DurationSecond` fuzzing [datafusion]
via GitHub
-
Re: [PR] Parse the DEFERRABLE transaction mode [datafusion-sqlparser-rs]
via GitHub
-
[PR] fix: fall back from the native Iceberg scan when a nested field was added or renamed [datafusion-comet]
via GitHub
-
Re: [PR] `PostgreSQL`: Support the `VARIADIC` function call argument marker [datafusion-sqlparser-rs]
via GitHub
-
Re: [PR] Added derive for arbitrary [datafusion-sqlparser-rs]
via GitHub
-
[PR] Add skewness metric to nodes that execute in partitioned mode [datafusion]
via GitHub
-
[I] Pre-select before cheap conjuncts when the undecided rows form a few long runs [datafusion]
via GitHub
-
[PR] chore: bump Ballista version to 55.0.0 [datafusion-ballista]
via GitHub
-
[PR] fix: dispatch StaticInvoke and Invoke only into Spark's own classes [datafusion-comet]
via GitHub
-
Re: [PR] fix: support struct-typed scalar subquery results [datafusion-comet]
via GitHub
-
Re: [PR] fix: preserve Spark errors for Parquet timestamp overflow [datafusion-comet]
via GitHub
-
Re: [PR] Add ExpressionAnalyzer for pluggable expression-level statistics estimation [datafusion]
via GitHub
-
Re: [PR] feat: support map inputs for explode [datafusion-comet]
via GitHub
-
Re: [I] Improve performance of first/last aggregates [datafusion-comet]
via GitHub
-
[PR] oss: add benches remaining scalars [datafusion-comet]
via GitHub
-
[PR] refactor: deprecate statistics providers that duplicate operator estimates [datafusion]
via GitHub
-
[PR] fix: [branch-1.1] defer Parquet conversion errors until a row group is decoded, as Spark does (#6515) [datafusion-comet]
via GitHub
-
[PR] fix: keep decimal fractions in modulo by one [datafusion]
via GitHub
-
Re: [I] Fuse operations in `equal_rows_arr` [datafusion]
via GitHub
-
Re: [PR] Add a fuzz target that flags superlinear parsing and printing [datafusion-sqlparser-rs]
via GitHub
-
[I] Final aggregate under AQE reads its shuffle through the JVM decoder instead of native direct read [datafusion-comet]
via GitHub
-
[PR] Js/store partitioning in dynamic filters [datafusion]
via GitHub
-
Re: [PR] fix: mark avg, bit_and/or/xor, stddev and variance as order-insensitive [datafusion]
via GitHub
-
Re: [I] Spark `xxhash64` ignores the running hash for dictionary-encoded values inside structs and lists [datafusion]
via GitHub
-
[PR] Show decimal add bug [datafusion]
via GitHub
-
Re: [PR] perf: bounded distinct count optimization [datafusion]
via GitHub
-
[PR] fix: hash every NaN like the canonical NaN in Spark xxhash64 [datafusion]
via GitHub
-
[PR] fix: array_concat panics or fails on non-default nested lists [datafusion]
via GitHub