costas-db opened a new issue, #3832: URL: https://github.com/apache/parquet-java/issues/3832
### Problem Per-column Hadoop configuration currently identifies a column with a flattened dot string such as `a.b`. This cannot distinguish two valid, different Parquet paths: ```text Top-level field named `a.b`: ["a.b"] Nested field `b` in `a`: ["a", "b"] ``` `ColumnConfigParser` passes the suffix to `ParquetProperties` as a string, and `ColumnProperty` interprets it with `ColumnPath.fromDotString`. Consequently, a setting intended for the top-level dotted field cannot be represented and may instead apply to the nested field. This is datatype-independent and affects per-column settings such as Bloom filters and statistics. ### Reproduction Use two ordinary string columns with the paths above, disable Bloom filters and statistics globally, and enable them only for the top-level `["a.b"]` path. On current `master`, the top-level column still has no Bloom filter or value statistics because the structured override is not recognized. ### Proposed direction Add a backward-compatible structured-path key form that encodes each path component independently, plus path-aware parser and builder overloads. Keep existing dot-string keys working for compatibility, while allowing callers to target dotted field names unambiguously. This changes only configuration lookup identity; it does not alter the Parquet file format. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
