ianmcook opened a new issue, #51779:
URL: https://github.com/apache/arrow/issues/51779

   > [!NOTE]
   > I discovered this issue and wrote it up with help from Claude Opus 5.5.
   
   ### Describe the bug, including details regarding any error messages, 
version, and platform.
   
   The format allows `DictionaryEncoding.indexType` to be omitted. 
[`format/Schema.fbs`](https://github.com/apache/arrow/blob/e85e181188ac3c5e8407b2160331faf5e78fd429/format/Schema.fbs#L500-L505)
 says:
   
   > If this field is null, the indices must be signed int32.
   
   The C++ IPC reader instead [requires the 
field](https://github.com/apache/arrow/blob/e85e181188ac3c5e8407b2160331faf5e78fd429/cpp/src/arrow/ipc/metadata_internal.cc#L899-L902):
   
   ```cpp
   auto int_data = encoding->indexType();
   CHECK_FLATBUFFERS_NOT_NULL(int_data, "DictionaryEncoding.indexType");
   ```
   
   So PyArrow can't read an Arrow IPC stream whose dictionary-encoded field 
omits `indexType`:
   
   ```
   OSError: Unexpected null field DictionaryEncoding.indexType in 
flatbuffer-encoded metadata
   ```
   
   #### Reproduction
   
   PyArrow and apache-arrow (JavaScript) both set `indexType` when they write, 
so this script embeds a stream without it. The reader fails on the schema, so 
the stream holds only a schema: one field, `text`, of type 
`dictionary<values=string, indices=int32>`, with `indexType` omitted. It was 
made from a stream that PyArrow wrote; the script that made it is below.
   
   ```python
   """Read an Arrow IPC stream whose DictionaryEncoding omits indexType."""
   
   import base64
   
   import pyarrow as pa
   
   # A schema-only stream with one field, text: dictionary-encoded strings, with
   # DictionaryEncoding.indexType omitted, so the indices are int32 per 
Schema.fbs.
   data = base64.b64decode(
       
"/////5gAAAAQAAAAAAAKAAwABgAFAAgACgAAAAABBAAEAAAAuP///wQAAAABAAAAFAAAABAAGAAI"
       
"AAYABwAMABAAFAAQAAAAAAABBRQAAABEAAAAIAAAAAQAAAAAAAAABAAAAHRleHQAAAAACAAIAAAA"
       
"BADc////DAAAAAgADAAIAAcACAAAAAAAAAEgAAAABAAEAAQAAAAIAAgAAAAAAP////8AAAAA"
   )
   print(pa.ipc.open_stream(data).schema)
   ```
   
   **Expected:** `text: dictionary<values=string, indices=int32, ordered=0>`
   
   **Actual:** `OSError: Unexpected null field DictionaryEncoding.indexType in 
flatbuffer-encoded metadata`
   
   apache-arrow (JavaScript) 21.2.0 reads the same stream as `Dictionary<Int32, 
Utf8>`, as the spec describes.
   
   The error is the same in PyArrow 12.0.1, 16.1.0, 20.0.0, 24.0.0, and 25.0.1 
(Python 3.11 and 3.14, macOS), so this isn't a regression.
   
   <details>
   <summary>How the stream was made</summary>
   
   This script writes a schema-only stream with PyArrow, then omits `indexType` 
from its schema message. With PyArrow 25.0.1 it prints exactly the base64 above.
   
   ```python
   """Write the stream above: a schema-only stream from PyArrow, with indexType 
omitted."""
   
   import base64
   import struct
   
   import pyarrow as pa
   
   
   def table_at(buf, pos):
       """Return the (table, vtable) positions for the table referenced at 
pos."""
       table = pos + struct.unpack_from("<I", buf, pos)[0]
       return table, table - struct.unpack_from("<i", buf, table)[0]
   
   
   def field_offset(buf, table, vtable, slot):
       """Return a field's offset within its table, or 0 if the field is 
absent."""
       entry = 4 + 2 * slot
       present = entry < struct.unpack_from("<H", buf, vtable)[0]
       return struct.unpack_from("<H", buf, vtable + entry)[0] if present else 0
   
   
   def without_index_type(stream):
       """Omit DictionaryEncoding.indexType from the first field of the schema 
message.
   
       Flatbuffers share identical vtables between tables (here the Schema's 
and the
       DictionaryEncoding's), so the encoding gets its own copy of its vtable, 
appended
       to the message, with indexType cleared.
       """
       buf = bytearray(stream)
       length = struct.unpack_from("<i", buf, 4)[0]  # After the 0xFFFFFFFF 
continuation marker
       meta = 8
       message, message_vt = table_at(buf, meta)
       schema, schema_vt = table_at(buf, message + field_offset(buf, message, 
message_vt, 2))  # Message.header
       fields = schema + field_offset(buf, schema, schema_vt, 1)  # 
Schema.fields
       field, field_vt = table_at(buf, fields + struct.unpack_from("<I", buf, 
fields)[0] + 4)  # fields[0]
       encoding, encoding_vt = table_at(buf, field + field_offset(buf, field, 
field_vt, 4))  # Field.dictionary
       vtable = bytearray(buf[encoding_vt:encoding_vt + 
struct.unpack_from("<H", buf, encoding_vt)[0]])
       struct.pack_into("<H", vtable, 4 + 2 * 1, 0)  # 
DictionaryEncoding.indexType: absent
       struct.pack_into("<i", buf, encoding, encoding - (meta + length))
       metadata = bytes(buf[meta:meta + length]) + bytes(vtable)
       metadata += bytes(-len(metadata) % 8)
       return struct.pack("<Ii", 0xFFFFFFFF, len(metadata)) + metadata + 
bytes(buf[meta + length:])
   
   
   schema = pa.schema({"text": pa.dictionary(pa.int32(), pa.string())})
   sink = pa.BufferOutputStream()
   with pa.ipc.new_stream(sink, schema):
       pass
   
print(base64.b64encode(without_index_type(sink.getvalue().to_pybytes())).decode())
   ```
   
   </details>
   
   #### Suggested fix
   
   Default to int32 when the field is absent:
   
   ```cpp
   std::shared_ptr<DataType> index_type;
   auto int_data = encoding->indexType();
   if (int_data == nullptr) {
     // Format/Schema.fbs: "If this field is null, the indices must be signed 
int32."
     index_type = int32();
   } else {
     RETURN_NOT_OK(IntFromFlatbuffer(int_data, &index_type));
   }
   ```
   
   #### Related
   
   - I found this through an apache-arrow (JavaScript) writer bug that drops 
`indexType`, among other problems; that's reported separately in as 
https://github.com/apache/arrow-js/issues/496. This issue is only about reading 
a stream that is valid under the spec.
   - nanoarrow crashes on a stream like this one; I will report that separately 
in apache/arrow-nanoarrow.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to