[ https://issues.apache.org/jira/browse/HADOOP-17112?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]
Steve Loughran resolved HADOOP-17112. ------------------------------------- Fix Version/s: 3.3.1 Resolution: Fixed fixed with test. Thank you for finding this! and providing the fix! > whitespace not allowed in paths when saving files to s3a via committer > ---------------------------------------------------------------------- > > Key: HADOOP-17112 > URL: https://issues.apache.org/jira/browse/HADOOP-17112 > Project: Hadoop Common > Issue Type: Sub-task > Components: fs/s3 > Affects Versions: 3.2.0 > Reporter: Krzysztof Adamski > Assignee: Krzysztof Adamski > Priority: Blocker > Labels: pull-request-available > Fix For: 3.3.1 > > Attachments: image-2020-07-03-16-08-52-340.png > > Time Spent: 1h > Remaining Estimate: 0h > > When saving results through spark dataframe on latest 3.0.1-snapshot compiled > against hadoop-3.2 with the following specs > --conf > spark.hadoop.mapreduce.outputcommitter.factory.scheme.s3a=org.apache.hadoop.fs.s3a.commit.S3ACommitterFactory > > --conf > spark.sql.parquet.output.committer.class=org.apache.spark.internal.io.cloud.BindingParquetOutputCommitter > > --conf > spark.sql.sources.commitProtocolClass=org.apache.spark.internal.io.cloud.PathOutputCommitProtocol > > --conf spark.hadoop.fs.s3a.committer.name=partitioned > --conf spark.hadoop.fs.s3a.committer.staging.conflict-mode=replace > we are unable to save the file with whitespace character in the path. It > works fine without. > I was looking into the recent commits with regards to qualifying the path, > but couldn't find anything obvious. Is this a known bug? > When saving results through spark dataframe on latest 3.0.1-snapshot compiled > against hadoop-3.2 with the following specs > --conf > spark.hadoop.mapreduce.outputcommitter.factory.scheme.s3a=org.apache.hadoop.fs.s3a.commit.S3ACommitterFactory > > --conf > spark.sql.parquet.output.committer.class=org.apache.spark.internal.io.cloud.BindingParquetOutputCommitter > > --conf > spark.sql.sources.commitProtocolClass=org.apache.spark.internal.io.cloud.PathOutputCommitProtocol > > --conf spark.hadoop.fs.s3a.committer.name=partitioned > --conf spark.hadoop.fs.s3a.committer.staging.conflict-mode=replace > we are unable to save the file with whitespace character in the path. It > works fine without. > I was looking into the recent commits with regards to qualifying the path, > but couldn't find anything obvious. Is this a known bug? > !image-2020-07-03-16-08-52-340.png! -- This message was sent by Atlassian Jira (v8.3.4#803005) --------------------------------------------------------------------- To unsubscribe, e-mail: common-dev-unsubscr...@hadoop.apache.org For additional commands, e-mail: common-dev-h...@hadoop.apache.org