behroozazarkhalili opened a new issue, #51305:
URL: https://github.com/apache/arrow/issues/51305

   
   Reproduced offline with **PyArrow 25.0.0**, Python 3.11.5, on Linux.
   The reproducer uses only PyArrow and Python's standard library; Hugging Face
   Datasets is not required. The downstream issue is huggingface/datasets#5531.
   
   ## Minimal reproducer
   
   ```python
   import io
   import json
   import pyarrow as pa
   import pyarrow.json as paj
   
   data = b'{"a":[null,1]}\n'  # 15 bytes, one valid JSON object
   assert json.loads(data) == {"a": [None, 1]}
   print(pa.__version__)
   table = paj.read_json(
       io.BytesIO(data),
       read_options=paj.ReadOptions(block_size=len(data)),
   )
   array = table["a"].chunk(0)
   print(array.offsets.to_pylist(), len(array.values))
   table.validate(full=True)
   ```
   
   Observed output (the read itself succeeds):
   
   ```text
   25.0.0
   [0, 2] 1
   pyarrow.lib.ArrowInvalid: Column 0: In chunk 0: Invalid: Length spanned by 
list offsets (2) larger than values array (length 1)
   ```
   
   Expected: `validate(full=True)` succeeds, offsets are `[0, 2]`, the child 
array
   contains `[None, 1]`, and the table contains one row `{"a": [None, 1]}`.
   The list offsets count the leading null, but the child array is one element 
short.
   Two leading nulls (`[null,null,1]`) produce offsets spanning three elements 
with
   a child array of length one.
   
   ## Block-dependent example
   
   ```python
   import io
   import pyarrow as pa
   import pyarrow.json as paj
   
   data = b'{"a":[1]}\n{"a":[null,2]}\n'  # 25 bytes
   for block_size in (13, len(data)):
       table = paj.read_json(io.BytesIO(data), 
read_options=paj.ReadOptions(block_size=block_size))
       try:
           table.validate(full=True)
           print(block_size, "valid")
       except pa.ArrowInvalid as error:
           print(block_size, error)
   ```
   
   Observed:
   
   ```text
   13 Column 0: In chunk 1: Invalid: Length spanned by list offsets (2) larger 
than values array (length 1)
   25 valid
   ```
   
   The preceding typed list prevents corruption when both records are in one 
block.
   When they are parsed in separate blocks, the later record can corrupt its 
block.
   The one-record reproducer above needs no internal block boundary or long 
string.
   
   ## Reduction of the original attachment
   
   The local `sanity_oscar_en.jsonl` from datasets#5531 contains **74,639,409 
bytes,
   10,000 rows**, each accepted by `json.loads`. Direct PyArrow reads measured:
   
   | `block_size` | Chunks | Full validation |
   | --- | --- | --- |
   | 1,048,576 | 72 | Invalid: column 2, chunk 2, struct child 3, offsets 5,651 
vs. 5,647 child values |
   | 16,777,216 | 5 | Valid |
   | 74,639,409 | 1 | Valid |
   
   Column 2 is `meta`; struct child 3 is `line_identifications`, a
   `list<struct<label: string, prob: double>>`. Of the 10,000 lists, 6,706 
contain a
   null and 945 begin with one. The first failing 1 MiB chunk starts at 
zero-based
   source row 336, whose list begins with four nulls, matching the four missing 
child
   values in that chunk.
   
   Row bisection (trying 1 MiB, maximum row length, and subset length block 
sizes as
   alignment changed) reduced the attachment to **original row 2**, saved as
   `min.jsonl`: **3,410 bytes, one row**. A nonempty row set cannot be smaller 
in row
   count. Its `meta.line_identifications` list has 22 elements, including 10 
nulls;
   positions 0–3 are null, and position 4 is the first typed struct
   `{"label":"en","prob":0.8105751276}`. Null positions are
   `[0,1,2,3,5,8,13,15,17,19]`.
   
   Both `block_size=2048` (record spans a boundary) and `block_size=4096` (one 
block)
   return an invalid table. At 4096, full validation reports:
   
   ```text
   Column 2: In chunk 0: Invalid: Struct child array #3 invalid: Invalid: 
Length spanned by list offsets (22) larger than values array (length 18)
   ```
   
   Removing `text` and all unrelated fields still fails. Removing just the four
   leading nulls makes this row validate; the later nulls can remain. The 
record's
   long text and metadata are therefore unnecessary for this reproducer.
   
   Additional single-block controls run on 25.0.0:
   
   - `[1,null]`, `[null]`, and `[]`: valid.
   - A null list followed by a typed list in another row: valid.
   - `[null]` followed by `[1]` in another row: valid.
   - `[null,"s"]` and `[null,[1]]`: invalid, one child value missing.
   - A list inside a struct containing `[null,{"label":"en","prob":1.0}]`:
     invalid, one child value missing.
   
   These results implicate null-to-typed list-value inference within a parser 
block.
   They do not establish a particular C++ implementation defect; no Arrow C++ 
build
   or internal debugger investigation was performed.
   
   ## Downstream effect and workaround limits
   
   `read_json` returns normally with invalid arrays, so consumers may fail 
later.
   Before the datasets fix, the original-row regression failed at full 
validation
   with the 22-versus-18 error above. Tests with typed and leading-null records 
in
   separate blocks also failed before the fix.
   
   A whole-file single-block retry validates for the complete attachment, but 
cannot
   repair a file beginning with the minimal leading-null list. Such inputs need 
a
   different parser, for example pandas' JSON decoder followed by an 
object-typed
   DataFrame and Arrow conversion. The datasets fallback uses the decoder with
   `precise_float=True`; object dtype avoids automatic date conversion and 
rounding
   of nullable integers through floats. Neither invalid tables nor their nested 
arrays should
   be converted to Python objects before successful validation.
   
   Only PyArrow 25.0.0 was tested in this investigation. Other Arrow versions,
   the gated dataset mentioned later in the issue, and the exact C++ cause 
remain
   unverified. This report has not been submitted upstream (offline task).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to