[
https://issues.apache.org/jira/browse/SPARK-18579?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15839704#comment-15839704
]
Hyukjin Kwon commented on SPARK-18579:
--------------------------------------
Can we just strip them within the dataframe/dataset? Available options for
read/write are well documented. IMHO, we should not just add options/APIs just
for consistency.
> spark-csv strips whitespace (pyspark)
> --------------------------------------
>
> Key: SPARK-18579
> URL: https://issues.apache.org/jira/browse/SPARK-18579
> Project: Spark
> Issue Type: Bug
> Components: Input/Output
> Affects Versions: 2.0.2
> Reporter: Adrian Bridgett
> Priority: Minor
>
> ignoreLeadingWhiteSpace and ignoreTrailingWhiteSpace are supported on CSV
> reader (and defaults to false).
> However these are not supported options on the CSV writer and so the library
> defaults take place which strips the whitespace.
> I think it would make the most sense if the writer semantics matched the
> reader (and did not alter the data) however this would be a change in
> behaviour. In any case it'd be great to have the _option_ to strip or not.
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]