Arawoof06 opened a new issue, #51641:
URL: https://github.com/apache/arrow/issues/51641
### Describe the bug, including details regarding any error messages,
version, and platform.
`DictDecoderImpl<ByteArrayType>::SetDict` and
`DictDecoderImpl<FLBAType>::SetDict` in `cpp/src/parquet/decoder.cc`
concatenate the decoded dictionary values into `byte_array_data_`, which is
resized to a 64-bit `total_size`, but they walk it with a 32-bit `offset`
accumulator:
```cpp
// ByteArray
int32_t offset = 0;
for (int i = 0; i < dictionary_length_; ++i) {
memcpy(bytes_data + offset, dict_values[i].ptr, dict_values[i].len);
...
offset += dict_values[i].len;
}
// FLBA
for (int32_t i = 0, offset = 0; i < dictionary_length_; ++i, offset +=
fixed_len) {
memcpy(bytes_data + offset, dict_values[i].ptr, fixed_len);
...
}
```
When the concatenated dictionary exceeds `INT32_MAX` bytes, `offset` wraps
negative and `memcpy(bytes_data + offset, ...)` writes outside the allocated
buffer (heap out-of-bounds write). The size is driven entirely by the
dictionary page of an untrusted Parquet file: a `DICTIONARY_PAGE` whose decoded
values total more than 2 GB reaches `SetDict` and triggers the wrap.
`total_size` is already computed as `int64_t`, so the accumulator was simply
left too narrow. For FLBA the concatenation is addressed directly by `index *
type_length` and can legitimately be larger than 2 GB, so the offset should be
64-bit. For `BYTE_ARRAY` the values are exposed through int32 offsets
(`byte_array_offsets_`), so a concatenation past `INT32_MAX` is unrepresentable
and should be rejected rather than silently wrapped.
### Component(s)
C++, Parquet
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]