ngoldbaum opened a new issue, #51760:
URL: https://github.com/apache/arrow/issues/51760
### Describe the enhancement requested
Casting a floating-point array to string writes integral values without a
decimal point:
```python
>>> import pyarrow as pa, pyarrow.compute as pc
>>> pc.cast(pa.array([20.0, 20.5, 1e10]), pa.string()).to_pylist()
['20', '20.5', '1e+10']
```
The CSV writer formats columns through the same cast, so float64 columns are
written the same way. Because the text no longer says the value was a float, a
reader that infers types (including `pyarrow.csv.read_csv`) reads the column
back as `int64`. Whether a column round-trips as `double` or `int64` then
depends on the values in it, not on its type.
#### Example
Writing a time series with one CSV per timestep, where the first timestep
happens to hold whole numbers:
```python
import pyarrow as pa, pyarrow.csv as csv, pyarrow.dataset as ds
csv.write_csv(pa.table({"x": [20.0, 21.0]}), "ts/t0.csv") # written as 20,
21
csv.write_csv(pa.table({"x": [20.5, 21.2]}), "ts/t1.csv")
ds.dataset("ts", format="csv").to_table()
# ArrowInvalid: Could not open CSV input source 'ts/t1.csv': Invalid: In CSV
column #0:
# Row #2: CSV conversion error to int64: invalid value '20.5'
```
Both tables are `float64`, but `t0.csv` is inferred as `int64`. pandas
`to_csv` and Polars `write_csv` both write `20.0`.
For context, this came up downstream in Narwhals, where `cast(String)` on a
float column gives different text on the PyArrow backend than with pandas or
Polars: https://github.com/narwhals-dev/narwhals/issues/4027
### Component(s)
C++, Python
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]