Arawoof06 opened a new issue, #51641:
URL: https://github.com/apache/arrow/issues/51641

   ### Describe the bug, including details regarding any error messages, 
version, and platform.
   
   `DictDecoderImpl<ByteArrayType>::SetDict` and 
`DictDecoderImpl<FLBAType>::SetDict` in `cpp/src/parquet/decoder.cc` 
concatenate the decoded dictionary values into `byte_array_data_`, which is 
resized to a 64-bit `total_size`, but they walk it with a 32-bit `offset` 
accumulator:
   
   ```cpp
   // ByteArray
   int32_t offset = 0;
   for (int i = 0; i < dictionary_length_; ++i) {
     memcpy(bytes_data + offset, dict_values[i].ptr, dict_values[i].len);
     ...
     offset += dict_values[i].len;
   }
   
   // FLBA
   for (int32_t i = 0, offset = 0; i < dictionary_length_; ++i, offset += 
fixed_len) {
     memcpy(bytes_data + offset, dict_values[i].ptr, fixed_len);
     ...
   }
   ```
   
   When the concatenated dictionary exceeds `INT32_MAX` bytes, `offset` wraps 
negative and `memcpy(bytes_data + offset, ...)` writes outside the allocated 
buffer (heap out-of-bounds write). The size is driven entirely by the 
dictionary page of an untrusted Parquet file: a `DICTIONARY_PAGE` whose decoded 
values total more than 2 GB reaches `SetDict` and triggers the wrap.
   
   `total_size` is already computed as `int64_t`, so the accumulator was simply 
left too narrow. For FLBA the concatenation is addressed directly by `index * 
type_length` and can legitimately be larger than 2 GB, so the offset should be 
64-bit. For `BYTE_ARRAY` the values are exposed through int32 offsets 
(`byte_array_offsets_`), so a concatenation past `INT32_MAX` is unrepresentable 
and should be rejected rather than silently wrapped.
   
   ### Component(s)
   
   C++, Parquet
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to