andygrove opened a new issue, #6321: URL: https://github.com/apache/datafusion-comet/issues/6321
Triage pass over the open `requires-triage` queue, per the project [Bug Triage Guide](https://github.com/apache/datafusion-comet/blob/main/docs/source/contributor-guide/bug_triage.md). - Date: 2026-09-28 - Total issues processed: 82 (80 triaged, 2 skipped, 0 failed) - Type counts: 46 bugs, 34 enhancements - Priority counts applied: `priority:critical` 15, `priority:high` 3, `priority:medium` 20, `priority:low` 8 - Guide: [docs/source/contributor-guide/bug_triage.md](https://github.com/apache/datafusion-comet/blob/main/docs/source/contributor-guide/bug_triage.md) Labels have already been applied. A reviewer should spot-check the calls below and close this issue when satisfied. Corrections should be made directly on the affected issue. Notes on this pass: - **Reviewer corrections from earlier passes are folded in.** After the 2026-09-14 and 2026-09-21 passes, reviewers raised #5801 and #5936 from `priority:medium` to `priority:critical`. In both, Comet returns rows where Spark raises an error. This pass applies that rule directly, so #6136 and #6086 are critical. Two other corrections are cited in the escalations below: a reviewer raised the metrics defect #5879 to high, and lowered #5701 from critical to medium. - **No type reclassifications.** Every `bug` or `enhancement` label an author had already applied matched the issue's content. The 20 issues that arrived without a type label were classified from their bodies. - **Two author priorities were escalated to critical:** #6288 (from `priority:high`) and #6290 (from `priority:medium`); see the escalations below. Author priorities were confirmed unchanged on #6203, #6217, #6283, #6289, #6292, #6294, #6315 and #6316. - **Opt-in features.** The native Iceberg writer (#6138 to #6146) and accelerated `mapInArrow` (#6290) are experimental and off by default. Silent wrong results there are labeled critical, as #5719 was, with an escalation note. The exception is #6145, which needs dates past year 262143 and is held at medium. The native Parquet writer (#6315) also needs an explicit `allowIncompatible` opt-in, so it follows #5131 and stays at medium. - #6291 is an enhancement that arrived with the author's `priority:low`. As in earlier passes, it was left in place, and the reviewer may want to remove it. - The guide's area table still lacks `area:memory`, `area:Iceberg`, `area:udf` and `area:joins`, so this pass did not add them. Where an issue already carries one, it is listed as "also carries" below, along with other labels outside the guide's table. - The guide lists `spark 4` as an area indicator, but the repository only has `spark 4.0` / `spark 4.1` / `spark 4.2` / `spark 3.x`. So nothing was applied to the Spark 4-only issues #6129, #6158, #6232, #6233 and #6317. - `regression` is not in the guide, so this pass neither added nor removed it. #6313 carries it, but the repository describes the label as "A bug that did not affect the most recent Comet release", while the issue dates the behavior to #3553 in 0.14.0. The reviewer may want to remove it there. #6225 does fit that description: #5174 introduced it after 1.0.0, and #5174 is in the 1.1 branch. - #6315, #6316 and #6317 were opened while this pass was running, and are included. ## Bugs ### priority:critical - array_append: ANSI item error is swallowed when the array is NULL ([#6086](https://github.com/apache/datafusion-comet/issues/6086)) - Area labels: `area:expressions` - Rationale: Under ANSI mode Comet returns NULL for a row where Spark raises `DIVIDE_BY_ZERO`, and `getSupportLevel` reports `Compatible`, so the error is lost with no fallback. That is step 1, consistent with #5218 and with the reviewer corrections that raised #5801 and #5936 to critical. Reachable on Spark 3.4 and 3.5 only. - Field id matching misses a container id after INT96 coercion drops container metadata ([#6131](https://github.com/apache/datafusion-comet/issues/6131)) - Area labels: `area:scan` - Rationale: With `spark.sql.parquet.fieldId.read.enabled=true`, the native scan returns nulls for an id-matched struct that contains a timestamp, where Spark returns the data. That is silent wrong results at step 1. The Spark setting is opt-in, but the INT96 timestamps that trigger it are Spark's default encoding. - CometNativeScanExec evaluates the inner (unrewritten) dynamic-pruning subquery from outputPartitioning: q5 fails with SubqueryAdaptiveBroadcastExec.execute(), q64 intermittently returns 0 rows at SF1000 ([#6133](https://github.com/apache/datafusion-comet/issues/6133)) - Area labels: `area:scan` - Rationale: User-reported on TPC-DS SF1000: q64 returns 0 rows where Spark returns 12,185, which is step 1 as the issue is titled. See the escalations for the split into #6264. - Nested duplicate names in a metadata-free Parquet file bypass the native resolver when the file schema equals the requested schema ([#6136](https://github.com/apache/datafusion-comet/issues/6136)) - Area labels: `area:scan` - Rationale: The native scan binds the colliding children by position and returns rows where Spark raises `foundDuplicateFieldInCaseInsensitiveModeError`. The 2026-09-14 pass put the duplicate-field-id sibling #5801 at medium and a reviewer raised it to critical, so this follows the corrected call. - Native Iceberg write merges -0.0 and 0.0 rows into one float/double identity partition ([#6138](https://github.com/apache/datafusion-comet/issues/6138)) - Area labels: `area:writer`; also carries `area:Iceberg`, `correctness` - Rationale: Rows are written under the other signed zero's partition value, so a filter that prunes on partition values can drop them with no error. That is the guide's data-corruption case at step 1, reachable only with the off-by-default native writer (see escalations). - Native Iceberg write silently drops S3 settings it cannot honour (credentials provider, SSE, ACL, tags, remote signing) instead of falling back ([#6139](https://github.com/apache/datafusion-comet/issues/6139)) - Area labels: `area:writer`; also carries `area:Iceberg`, `correctness` - Rationale: Objects can be written without the table owner's SSE-KMS or SSE-C encryption, ACL, tags or storage class, and nothing reports it. The guide's `priority:critical` row covers security vulnerabilities. The native writer is off by default (see escalations). - Comparison operators on arrays and structs with floating-point leaves do not match Spark for signed zero ([#6157](https://github.com/apache/datafusion-comet/issues/6157)) - Area labels: `area:expressions`; also carries `correctness` - Rationale: `<=>`, `<` and `>=` over nested `-0.0` and `0.0` run fully native and return the opposite answers from Spark, which is step 1. This is the same class as #6019, which the last pass put at critical. - CometSort sorts a multi-column key containing a collated string by raw bytes ([#6158](https://github.com/apache/datafusion-comet/issues/6158)) - Area labels: none; also carries `correctness` - Rationale: `row_number()` over a multi-column key with a `UTF8_LCASE` string comes back in byte order, with no fallback. That is silent wrong results at step 1, on Spark 4.0+ collations. - Codegen dispatcher writes a null map key as the key type's default value ([#6172](https://github.com/apache/datafusion-comet/issues/6172)) - Area labels: `area:expressions` - Rationale: On the default configuration, `map_keys(transform_values(try_cast(...)))` returns `[1, 0]` where Spark returns `[1, NULL]`, which is silent wrong results at step 1. - AQE reuses one exchange for scans with different dynamic pruning filters and drops rows ([#6264](https://github.com/apache/datafusion-comet/issues/6264)) - Area labels: `area:scan` - Rationale: A `ReusedExchange` hands one branch the other branch's rows. The reproduction returns 7 or 21 rows where Spark returns 28, with no error, which is step 1. This is the likely cause of the q64 loss in #6133. - Storage-partitioned self-join of an Iceberg table returns duplicate rows with partially clustered distribution ([#6278](https://github.com/apache/datafusion-comet/issues/6278)) - Area labels: `area:scan` - Rationale: On the default-on native Iceberg scan, Comet returns 378 rows where Spark returns 126, which is silent wrong results at step 1. - Sliced booleans nested in structs, and sliced inputs to the JVM UDF bridge, reach the JVM misaligned and return wrong results ([#6288](https://github.com/apache/datafusion-comet/issues/6288)) - Area labels: `area:ffi`; also carries `correctness` - Rationale: On default configs, after `LIMIT ... OFFSET` or a grouped aggregate with more than one batch of groups, Comet returns wrong boolean values and puts nulls on the wrong rows. The guide lists "boolean arrays with non-zero offset" at the FFI boundary as a critical example. Escalated from the author's `priority:high`. - Accelerated mapInArrow reads Python output at the declared types without checking them ([#6290](https://github.com/apache/datafusion-comet/issues/6290)) - Area labels: none; also carries `area:udf`, `correctness` - Rationale: Where Spark raises on an output type mismatch, the accelerated path reads the buffers at the declared width. It returns wrong integers or decimals, or reads past the end of a buffer, which is step 1. Escalated from the author's `priority:medium`; the feature is experimental and off by default (see escalations). - Nested DATE to numeric casts in structs and maps return the day count or fail ([#6316](https://github.com/apache/datafusion-comet/issues/6316)) - Area labels: `area:expressions`; also carries `correctness` - Rationale: In legacy mode, the Spark 3.x default, a `DATE` to `INT` cast inside a struct or map returns the day count where Spark returns NULL. The plan is fully native and there is no error, which is step 1; the other numeric and boolean targets fail the query with an internal error. Confirms the author's label. - A native plan with no JVM input returns truncated output without an error when the Tokio runtime shuts down ([#6294](https://github.com/apache/datafusion-comet/issues/6294)) - Area labels: `area:ffi`; also carries `correctness` - Rationale: When an executor shuts down mid-task, the task reports success with truncated output: 1.2M of 4.8M rows in the reproduction. That is silent wrong results at step 1, and confirms the author's label. ### priority:high - arrays_zip with two same-named inputs fails with "ArrowArray struct has 2 children (expected 1)" ([#6251](https://github.com/apache/datafusion-comet/issues/6251)) - Area labels: `area:expressions`, `area:ffi` - Rationale: On default configs the task dies with an `IllegalStateException` in the Arrow import instead of falling back. Two same-named inputs arise naturally after a join. That is an unhandled exception on a supported path, which the guide rates `priority:high`, as with #5059. - JVM columnar shuffle fails with ArrayIndexOutOfBoundsException when spark.shuffle.checksum.enabled=false ([#6256](https://github.com/apache/datafusion-comet/issues/6256)) - Area labels: `area:shuffle` - Rationale: With Spark's shuffle checksums disabled, every task of a sort-based JVM columnar shuffle throws an unhandled `ArrayIndexOutOfBoundsException`. That matches the guide's "NPE on supported code path" example at step 2, and it affects 1.0.0 and 1.1.0. - spark.comet.batchSize below 8192 makes CometConf fail to initialize on every executor ([#6286](https://github.com/apache/datafusion-comet/issues/6286)) - Area labels: none - Rationale: A valid tuning value makes `CometConf` throw `ExceptionInInitializerError`, then `NoClassDefFoundError` for every later task on the executor, so the job aborts. That is an unhandled exception that breaks every query, step 2. ### priority:medium - Iceberg scan fails with a hard-wired 10s OpenDAL io_timeout that cannot be configured ([#6124](https://github.com/apache/datafusion-comet/issues/6124)) - Area labels: `area:scan`; also carries `area:Iceberg` - Rationale: A native Iceberg read fails on a 10-second timeout that iceberg-java does not impose. The failure is visible and there are workarounds: disable the native Iceberg scan, or raise `COMET_WORKER_THREADS` if the cause is worker starvation. Step 3. - Native Iceberg write gate reads hdfs:/path as file, so the write fails natively instead of falling back ([#6140](https://github.com/apache/datafusion-comet/issues/6140)) - Area labels: `area:writer`; also carries `area:Iceberg` - Rationale: The write fails with `Unsupported storage scheme: hdfs` instead of falling back. The failure is visible and needs the off-by-default native writer, so step 3. The default-on scan analog, #5541, is `priority:high`. - Native Iceberg write fails on a V1 spec that mixes a live field with a void field whose source column was dropped ([#6141](https://github.com/apache/datafusion-comet/issues/6141)) - Area labels: `area:writer`; also carries `area:Iceberg` - Rationale: Every task fails where iceberg-java succeeds, with no fallback. The failure is visible and needs the off-by-default native writer, so step 3. - Native Iceberg write panics or writes a NULL partition value for dates and timestamps beyond year 262143 ([#6145](https://github.com/apache/datafusion-comet/issues/6145)) - Area labels: `area:writer`; also carries `area:Iceberg` - Rationale: The NULL partition value is silent wrong metadata, but it needs dates past year 262143 on the off-by-default writer. Held at medium the way #5256 was held for a correctness path that cannot be reached in practice (see escalations). The author called it low priority. - Native Iceberg write over-counts NaNs for nested float fields when the input batch is already sliced ([#6146](https://github.com/apache/datafusion-comet/issues/6146)) - Area labels: `area:writer`; also carries `area:Iceberg` - Rationale: Data files get wrong `nan_value_counts` for list and map float elements, which no predicate can prune on. So this is broken metadata rather than wrong query results, step 3. - RevertNativeForTransitionHeavyStages strips transitions in the stage below when AQE is off ([#6152](https://github.com/apache/datafusion-comet/issues/6152)) - Area labels: none - Rationale: Found by code reading: the rule can leave a row operator over a columnar child. That needs both the off-by-default `spark.comet.exec.transitionRevert.enabled` and AQE off, so step 3. - `CometCast.isSupported` returns the first non-Compatible child, so `Incompatible` can mask `Unsupported` ([#6200](https://github.com/apache/datafusion-comet/issues/6200)) - Area labels: `area:expressions` - Rationale: For a struct or map cast with an `Unsupported` child, the support-level check reports `Incompatible`, and under `allowIncompatible` that child would run natively. No query reaches it today, because the only `Incompatible` cast needs a negative-scale decimal that cannot normally be built, so it is held at medium (see escalations). - ANSI integral overflow errors: wrong error class for Byte/Short arithmetic and missing try_ suggestion ([#6217](https://github.com/apache/datafusion-comet/issues/6217)) - Area labels: `area:expressions` - Rationale: Only the error class and message differ; the decision to throw matches Spark. That is the same tier as its predecessor #5071 and as #5073, and confirms the author's label. - fair_unified: a JVM consumer freeing its last bytes can fail a parked native acquire with NoSuchElementException ([#6224](https://github.com/apache/datafusion-comet/issues/6224)) - Area labels: none; also carries `area:memory` - Rationale: A race in Spark's `ExecutionMemoryPool` fails the task with `NoSuchElementException` instead of refusing the request so the operator can spill. It is shown at component level only, and the task fails visibly, so step 3. - Reduce retained buffer allocation when list_extract selects short nested arrays ([#6225](https://github.com/apache/datafusion-comet/issues/6225)) - Area labels: `area:expressions` - Rationale: #5174, which is in the 1.1 branch, makes `list_extract` retain about 32 times the buffer capacity when it selects short inner arrays, and native shuffle charges that capacity against its reservation. The values are correct and there is no end-to-end measurement yet, so this is a memory regression at step 3. - Spark errors after the first batch of a JVM input stream reach the user as CometNativeException ([#6234](https://github.com/apache/datafusion-comet/issues/6234)) - Area labels: `area:ffi` - Rationale: The query fails as it would in Spark, but the exception class and error condition depend on which batch failed. That is error fidelity, the same tier as #5073. - Grouped integer SUM doesn't count its per-group state in the aggregate's memory reservation ([#6252](https://github.com/apache/datafusion-comet/issues/6252)) - Area labels: `area:aggregation`; also carries `area:memory` - Rationale: `size()` reports the struct rather than its per-group vectors, so every grouped integer `SUM` under-reserves by 16 bytes per group and spills late. The results are correct, so step 3. - Native window operators reserve no memory for the batches they buffer ([#6253](https://github.com/apache/datafusion-comet/issues/6253)) - Area labels: none; also carries `area:memory` - Rationale: `WindowAggExec` runs Spark's default frame for `agg(...) OVER (PARTITION BY k)`. It keeps every input batch of the task's partition with no reservation, where Spark's `WindowExec` spills, so a large partition can push the executor past its container limit. Not yet measured, so step 3. - A native final aggregate that has spilled can fail the task during its replay ([#6254](https://github.com/apache/datafusion-comet/issues/6254)) - Area labels: `area:aggregation`; also carries `area:memory` - Rationale: The replay after a spill fails the task with `Additional allocation failed` instead of spilling again, and it reproduces. The failure is visible and falling back to Spark's aggregate avoids it, so step 3; see the escalations. - Hash-based JVM columnar shuffle reports its output as disk spill instead of bytes written ([#6258](https://github.com/apache/datafusion-comet/issues/6258)) - Area labels: `area:shuffle` - Rationale: Shuffle bytes written are under-reported and "Spill (Disk)" shows most of the output, while `MapStatus` stays correct. That is a defect in existing metrics reporting, classified like #5336 and #5382 (see escalations). - Spark 3.4 and 3.5 release jars require Java 17 since 0.11.0, though the docs list Java 11 ([#6283](https://github.com/apache/datafusion-comet/issues/6283)) - Area labels: `area:ci`; also carries `build` - Rationale: The published Spark 3.x jars fail to load on Java 11 with `UnsupportedClassVersionError`. The failure is visible and running on Java 17 avoids it, and the cause is in the release build tooling. Confirms the author's label. - Comet starts a single Tokio worker on standalone executors when spark.executor.cores is unset ([#6292](https://github.com/apache/datafusion-comet/issues/6292)) - Area labels: none; also carries `performance` - Rationale: A standalone executor with `spark.executor.cores` unset, the standalone default, gets a single Tokio worker. In the reporter's measurement, a native scan feeding a sort ran 2.4 times slower with one worker than with four, which is significant performance degradation at step 3. Confirms the author's label; see the escalations for the deadlock. - A JVM consumer's parked page allocation fails when another consumer of the task empties its balance ([#6304](https://github.com/apache/datafusion-comet/issues/6304)) - Area labels: `area:shuffle` - Rationale: This is the JVM-side twin of #6224: a parked `allocatePage` in `CometUnifiedShuffleMemoryAllocator` fails with `NoSuchElementException`. It is shown at component level only, and the task fails visibly, so step 3. - Native Parquet writes on Spark 3.4/3.5 leave INSERT INTO targets stale until REFRESH TABLE ([#6315](https://github.com/apache/datafusion-comet/issues/6315)) - Area labels: `area:writer` - Rationale: On Spark 3.4 and 3.5, a native `INSERT INTO ... SELECT` reads back empty until the table is refreshed, which is silent. But the native Parquet write needs both the Testing-category `spark.comet.parquet.write.enabled` and the `DataWritingCommandExec` `allowIncompatible` opt-in, so as with #5131, users only see it after accepting incompatible behavior. Confirms the author's label (see escalations). - Native blocks without a JVM input publish SQL metrics on every batch, ignoring spark.comet.metrics.updateInterval ([#6313](https://github.com/apache/datafusion-comet/issues/6313)) - Area labels: none; also carries `performance`, `regression` - Rationale: Native blocks with no JVM input ignore `spark.comet.metrics.updateInterval` and publish metrics on every batch. A one-file scan took 343 ms of process CPU against 209 ms with the interval honored. Performance degradation since #3553 (0.14.0), step 3. ### priority:low - native: `JVMClasses::with_env: JAVA_VM not initialized` aborts a multi-suite JVM ([#6096](https://github.com/apache/datafusion-comet/issues/6096)) - Area labels: `area:ffi` - Rationale: The JNI lifecycle abort has only been seen when 132 Comet suites share one JVM with debug assertions on, and CI splits the suites and passes. That is a test-only failure at step 4 (see escalations). - Three dev/diffs weaken a local-shuffle-read assertion into a contradiction, breaking the Spark-only baseline ([#6122](https://github.com/apache/datafusion-comet/issues/6122)) - Area labels: `spark sql tests` - Rationale: The patched test cannot pass for any number of local reads. It is skipped under Comet and breaks only the Spark-only baseline run, so this is test-only, step 4. - Native memory usage log underestimates the executor overhead for PySpark and SparkR on Kubernetes, and warns on standalone clusters ([#6188](https://github.com/apache/datafusion-comet/issues/6188)) - Area labels: none; also carries `area:memory` - Rationale: The log warns of a container kill that cannot happen, and warns on standalone clusters that have no container. That is a misleading log line, step 4. - In-memory cache tests that use checkSparkAnswer compare the cache with itself ([#6203](https://github.com/apache/datafusion-comet/issues/6203)) - Area labels: none; also carries `test` - Rationale: 23 cache tests cannot catch a wrong stored value or a wrong pruning bound, which is a test-only problem at step 4. Confirms the author's label; it bears on the default-on decision in #5634. - With several native plans in one task, the non-zero memory usage warning comes from the wrong plan ([#6255](https://github.com/apache/datafusion-comet/issues/6255)) - Area labels: none; also carries `area:memory` - Rationale: The first plan logs a leak that is not one, and a real leak in a later plan is never reported. Log diagnostics only, step 4. - spark.comet.shuffle.jvm.batchSize=0 hangs a task, and an unknown spark.comet.exec.memoryPool fails every task ([#6259](https://github.com/apache/datafusion-comet/issues/6259)) - Area labels: `area:shuffle`; also carries `area:memory` - Rationale: Two settings are not validated, and the failures need an invalid value: a zero batch size, or a misspelled pool name that fails every task with a clear message. Step 4. - The native memory usage log understates untracked memory while a pool is overcommitted ([#6260](https://github.com/apache/datafusion-comet/issues/6260)) - Area labels: none; also carries `area:memory` - Rationale: While a pool is overcommitted, the log's untracked figure and the container warning come out low by the overcommit. The author notes that is usually small and short-lived. Log diagnostics, step 4. - An input ArrowArrayStream that native never takes is never released ([#6289](https://github.com/apache/datafusion-comet/issues/6289)) - Area labels: `area:ffi` - Rationale: The stream, and the first input batch it buffered, stay pinned for the executor's life only when a task ends before native takes the stream, for example a failed `createPlan` or an early cancellation. That is a leak on error paths, and confirms the author's label. ## Enhancements - Support Spark's `AtLeastNNonNulls` natively ([#6093](https://github.com/apache/datafusion-comet/issues/6093)) - Area labels: `area:expressions` - Rationale: New native expression support. Routing through the codegen dispatcher works but was measured slower than falling back. - ci: key the TPC dataset caches on pinned generators ([#6102](https://github.com/apache/datafusion-comet/issues/6102)) - Area labels: `area:ci` - Rationale: CI cache efficiency; no job fails today. - ci: shard the Iceberg extensions test task ([#6103](https://github.com/apache/datafusion-comet/issues/6103)) - Area labels: `area:ci` - Rationale: Shortens the longest job in the merge queue; no job fails today. - Iceberg S3-family storage rebuilds the opendal Operator and signer on every file open ([#6109](https://github.com/apache/datafusion-comet/issues/6109)) - Area labels: `area:scan` - Rationale: A performance optimization that needs an upstream iceberg-rust change. Reads are correct today. - Native Iceberg writer keeps a dictionary page for high-cardinality columns where iceberg-java writes none ([#6114](https://github.com/apache/datafusion-comet/issues/6114)) - Area labels: `area:writer`; also carries `area:Iceberg`, `performance` - Rationale: The issue states that results are correct either way. This is a file layout and read-performance improvement. - Charge native write buffers to the shared off-heap memory pool ([#6115](https://github.com/apache/datafusion-comet/issues/6115)) - Area labels: `area:writer`; also carries `area:Iceberg`, `area:memory` - Rationale: Charges bounded writer buffers to the memory pool, an accounting improvement classified like #6009 and #6063. - docs: add a user-facing Celeborn integration guide ([#6118](https://github.com/apache/datafusion-comet/issues/6118)) - Area labels: `area:shuffle`; also carries `documentation` - Rationale: Documentation addition. - Write blog post for 1.1.0 release ([#6120](https://github.com/apache/datafusion-comet/issues/6120)) - Area labels: none - Rationale: Release communication; no code change. - perf: reduce Parquet runtime filter schema guard overhead ([#6123](https://github.com/apache/datafusion-comet/issues/6123)) - Area labels: `area:scan` - Rationale: Wins back the pruning and per-file cost given up by the correctness guard in #6067, which its review accepted as follow-up work. This is a performance optimization of correct behavior. - Investigate charging native Iceberg scan memory to the task memory pool ([#6126](https://github.com/apache/datafusion-comet/issues/6126)) - Area labels: `area:scan`; also carries `area:Iceberg`, `area:memory` - Rationale: An investigation into extending #6125's scan memory accounting to the Iceberg scan. - Support scalar PySpark Arrow UDFs in Comet native execution via PyO3 ([#6129](https://github.com/apache/datafusion-comet/issues/6129)) - Area labels: none - Rationale: A new opt-in native operator for Spark 4.1+ `ArrowEvalPythonExec`, which falls back correctly today. - Report resident memory (RssAnon) in the executor memory usage log and warn on it against the container size ([#6167](https://github.com/apache/datafusion-comet/issues/6167)) - Area labels: none; also carries `area:memory` - Rationale: Observability addition to the memory usage log. - Follow-ups from the 1.1.0 user guide review ([#6169](https://github.com/apache/datafusion-comet/issues/6169)) - Area labels: none - Rationale: A checklist of documentation corrections and the checks behind them. Its two wrong-results items were fixed elsewhere: #5783 by #5786, and the Iceberg transform residuals by #6154. - Native UDFs: test coverage for empty batches and varied array encodings ([#6174](https://github.com/apache/datafusion-comet/issues/6174)) - Area labels: `area:ffi`; also carries `area:udf` - Rationale: Test coverage for C Data Interface edge cases; no failure is reported. - Native UDFs: config to disable or restrict library loading ([#6175](https://github.com/apache/datafusion-comet/issues/6175)) - Area labels: none; also carries `area:udf` - Rationale: A new config for operators to disable or restrict native UDF library loading. - Native UDFs: distribute the library to executors ([#6176](https://github.com/apache/datafusion-comet/issues/6176)) - Area labels: none; also carries `area:udf` - Rationale: New capability to ship the native UDF library to executors. - Native UDFs: lift the 4-argument cap and add `registerAll` ([#6177](https://github.com/apache/datafusion-comet/issues/6177)) - Area labels: none; also carries `area:udf` - Rationale: Extends the registration API: more than four arguments, and `registerAll`. - Follow-ups to #6163: correct the `fair_unified` description and prepare `spark.comet.exec.memoryPool.fraction` for removal ([#6187](https://github.com/apache/datafusion-comet/issues/6187)) - Area labels: none; also carries `documentation`, `area:memory` - Rationale: Documentation corrections and the remaining work for the config deprecation. - Kubernetes guide example never enables Comet because it sets no off-heap memory ([#6190](https://github.com/apache/datafusion-comet/issues/6190)) - Area labels: none; also carries `documentation` - Rationale: A documentation correction, which the guide's type table places under `enhancement`. - `ExtractANSIIntervalDays` and the other interval field extractors have no serde, so `date + <day interval column>` and `extract` of an interval fall back ([#6193](https://github.com/apache/datafusion-comet/issues/6193)) - Area labels: `area:expressions` - Rationale: New expression support; the fallback is correct today. - Charge the native shuffle writer's write buffers to the memory pool ([#6196](https://github.com/apache/datafusion-comet/issues/6196)) - Area labels: `area:shuffle`; also carries `area:memory` - Rationale: Accounting improvement for bounded buffers, about 4 MiB per task. Same class as #6115. - Release builds of libcomet skip LTO because the crate is also an rlib ([#6210](https://github.com/apache/datafusion-comet/issues/6210)) - Area labels: `area:ci`; also carries `build`, `performance` - Rationale: A build configuration change: measure LTO for `libcomet`, then possibly enable it. - perf: account native allocations without a thread-local lookup ([#6213](https://github.com/apache/datafusion-comet/issues/6213)) - Area labels: none; also carries `performance`, `area:memory` - Rationale: Performance optimization of the allocation accounting wrapper. - Follow-ups for the IRSA web-identity credential provider ([#6231](https://github.com/apache/datafusion-comet/issues/6231)) - Area labels: `area:scan`; also carries `area:Iceberg` - Rationale: Review follow-ups for #6025, which is still open, so none of this is on main yet. - `CometLiteral` accepts a collated string literal and serializes it as a plain string ([#6232](https://github.com/apache/datafusion-comet/issues/6232)) - Area labels: `area:expressions` - Rationale: The reporter confirms the answer is right today and asks for the behavior to be made explicit and tested. That is hardening, as with #5943. - Move `CometCollationSuite` to `spark-4.x` so every 4.x profile runs it ([#6233](https://github.com/apache/datafusion-comet/issues/6233)) - Area labels: none - Rationale: Moves the collation suite so that the Spark 4.2 profile runs it too. - Support Native Iceberg MOR write ([#6240](https://github.com/apache/datafusion-comet/issues/6240)) - Area labels: `area:writer` - Rationale: New native write capability. - CometTaskMemoryManager logs a warning and a memory dump every time a native reservation is refused ([#6257](https://github.com/apache/datafusion-comet/issues/6257)) - Area labels: none; also carries `area:memory` - Rationale: Changes the log level of a routine event; no behavior change. - Remove AlignedArrowStreamReader and fix the Native to JVM section of ffi.md ([#6291](https://github.com/apache/datafusion-comet/issues/6291)) - Area labels: `area:ffi`; also carries `documentation` - Rationale: Dead-code removal and a documentation rewrite, with no behavior change intended. - Blocking JVM calls from native plans hold Tokio workers, delaying other plans and I/O ([#6293](https://github.com/apache/datafusion-comet/issues/6293)) - Area labels: `area:ffi`; also carries `performance` - Rationale: A scheduling improvement for blocking JVM calls. The deadlock it relates to is tracked as the bug #6292. - Cancelling a task doesn't stop its native plan until the plan produces its next batch ([#6295](https://github.com/apache/datafusion-comet/issues/6295)) - Area labels: `area:ffi`; also carries `performance` - Rationale: Adds native cancellation so that a killed task frees its slot sooner. Results are unaffected, and the author filed it as an enhancement. - S3 credential refreshes are not coalesced, and bucket region detection has no timeout ([#6296](https://github.com/apache/datafusion-comet/issues/6296)) - Area labels: `area:scan` - Rationale: Hardening found by code reading, with no observed failure, as with #5943. - Support anonymous S3 access through the CometS3CredentialProvider SPI ([#6298](https://github.com/apache/datafusion-comet/issues/6298)) - Area labels: `area:scan` - Rationale: New capability for the credential provider SPI. Today the adapter fails with a clear error for anonymous buckets. - Support Spark 4.2 geometry type + functions ([#6317](https://github.com/apache/datafusion-comet/issues/6317)) - Area labels: `area:expressions` - Rationale: New type and function support for Spark 4.2's `GEOMETRY` and its initial `ST_*` functions. ## Escalations to consider - Sliced booleans nested in structs, and sliced inputs to the JVM UDF bridge, reach the JVM misaligned and return wrong results ([#6288](https://github.com/apache/datafusion-comet/issues/6288)) - Escalated in this pass from the author's `priority:high` to `priority:critical`. The guide's list of correctness bugs includes "Data corruption in FFI boundary (e.g., boolean arrays with non-zero offset)". The reproductions return wrong values on default configs through ordinary `LIMIT ... OFFSET` and grouped-aggregate plans. - Accelerated mapInArrow reads Python output at the declared types without checking them ([#6290](https://github.com/apache/datafusion-comet/issues/6290)) - Escalated in this pass from the author's `priority:medium` to `priority:critical` on decision-tree step 1. The feature is experimental and off by default (`spark.comet.exec.pyarrowUDF.enabled`), and the report comes from reading the code rather than a run. The reviewer may prefer to restore `priority:medium` under the "core path over experimental" principle. - array_append: ANSI item error is swallowed when the array is NULL ([#6086](https://github.com/apache/datafusion-comet/issues/6086)) - Labeled critical. Spark's interpreted `ArrayAppend.eval` short-circuits exactly as Comet does; only Spark's generated code raises. If the reviewer reads that as Spark disagreeing with itself, `priority:medium` fits, just as a reviewer lowered #5701 from critical to medium. - CometNativeScanExec evaluates the inner (unrewritten) dynamic-pruning subquery from outputPartitioning: q5 fails with SubqueryAdaptiveBroadcastExec.execute(), q64 intermittently returns 0 rows at SF1000 ([#6133](https://github.com/apache/datafusion-comet/issues/6133)) - Labeled critical as titled. The q64 row loss has since been split out as #6264 (fix PR #6270), and the q5 failure has its own fix PR, #6268. The reporter's SF1000 runs with both PRs applied match Spark. If #6264 is accepted as the q64 cause, what remains here is a `SparkUnsupportedOperationException` on the default scan selection, which is `priority:high` at step 2. - Native Iceberg write merges -0.0 and 0.0 rows into one float/double identity partition ([#6138](https://github.com/apache/datafusion-comet/issues/6138)) - Labeled critical, but reachable only with `spark.comet.iceberg.write.enabled=true`, a Testing-category setting that defaults to false. #5719 was labeled critical on the same basis and has kept it, but the reviewer may prefer `priority:high` under "core path over experimental". Either way, it should be settled before the native writer is turned on by default (#5644). - Native Iceberg write silently drops S3 settings it cannot honour (credentials provider, SSE, ACL, tags, remote signing) instead of falling back ([#6139](https://github.com/apache/datafusion-comet/issues/6139)) - Same off-by-default consideration as #6138. The critical label rests on the security side: the table owner's encryption, ACL and tag settings are dropped from written objects with no signal. The GCS counterpart #5637 was labeled `priority:high` because it was framed as a storage identity problem. - Native Iceberg write panics or writes a NULL partition value for dates and timestamps beyond year 262143 ([#6145](https://github.com/apache/datafusion-comet/issues/6145)) - Held at `priority:medium`. The `year` and `month` NULL partition value is silent wrong metadata, which step 1 would place at critical. But it needs dates beyond year 262143 on the off-by-default writer, so it is held at medium, as #5256 was. The author's own "low priority" reading is also defensible. - `CometCast.isSupported` returns the first non-Compatible child, so `Incompatible` can mask `Unsupported` ([#6200](https://github.com/apache/datafusion-comet/issues/6200)) - Held at `priority:medium` because no query reaches it today. If #6179 or a new `Incompatible` cast makes the masked `Unsupported` child reachable, that cast would run natively under `allowIncompatible`, with a crash (as in #5995) or a wrong result. Re-evaluate at that point. - A native final aggregate that has spilled can fail the task during its replay ([#6254](https://github.com/apache/datafusion-comet/issues/6254)) - Matches the guide's trigger "A `priority:medium` bug is reported by multiple users or affects a common workload → consider escalating to `priority:high`". Any grouped aggregate whose final stage spills under memory pressure can hit it. The upstream report, apache/datafusion#25423, was closed by a test-only change, so the behavior is unchanged in DataFusion 55.1. - Comet starts a single Tokio worker on standalone executors when spark.executor.cores is unset ([#6292](https://github.com/apache/datafusion-comet/issues/6292)) - Matches the same trigger. Leaving `spark.executor.cores` unset is the standalone default. On `main`, the single worker also deadlocked in 4 of 4 runs with 96 MB or 128 MB of off-heap memory. #6261, still open, fixes the deadlock but not the slowdown. - Hash-based JVM columnar shuffle reports its output as disk spill instead of bytes written ([#6258](https://github.com/apache/datafusion-comet/issues/6258)) - Kept at `priority:medium` alongside #5336 and #5382. However, on 2026-09-14 a reviewer raised a similar metrics under-report, #5879, to `priority:high`. - Native Parquet writes on Spark 3.4/3.5 leave INSERT INTO targets stale until REFRESH TABLE ([#6315](https://github.com/apache/datafusion-comet/issues/6315)) - Held at `priority:medium` because the native Parquet write sits behind an `allowIncompatible` opt-in, following #5131. The stale read is silent, so step 1 would place it at critical once the native Parquet writer runs without that opt-in. - native: `JVMClasses::with_env: JAVA_VM not initialized` aborts a multi-suite JVM ([#6096](https://github.com/apache/datafusion-comet/issues/6096)) - Labeled `priority:low` as a test-harness failure. If the `JAVA_VM` lifecycle gap can be hit by a long-lived production process, for example one that loads the native library a second time, the result is a JVM abort, which is `priority:high` at step 2. ## Skipped (needs more info) Both are tracking or record issues rather than bug reports or feature requests. No `bug` or `enhancement` label fits, so `requires-triage` was left in place, and they will reappear in each pass until they are closed. - [EPIC] Bug fixes to consider backporting to branch-1.0 ([#6201](https://github.com/apache/datafusion-comet/issues/6201)) - A release-management tracker for 1.0.1 backports, neither a defect nor a feature request. It carries forward #5815, which earlier passes skipped on the same grounds and which was closed when this one was opened. - Bug triage results: 2026-08-24 ([#5454](https://github.com/apache/datafusion-comet/issues/5454)) - The summary issue from the 2026-08-24 pass. None of its labels were applied at the time, but all 28 issues it covers have since been triaged: each carries a type label and none still carries `requires-triage`. The three earlier summaries it links are closed. Nothing in it is outstanding, so it can be closed. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
