sunchao opened a new issue, #3747:
URL: https://github.com/apache/parquet-java/issues/3747

   ## Describe the bug
   
   Column-index filtering can make the record-based Parquet reader return a 
value from the wrong row. The problem occurs when the reader takes the final 
selected row position from its iterator while it is still processing an earlier 
data page. It treats the now-empty iterator as completion, even though that 
last row has not been reached.
   
   Here is a concrete example using an ordinary equality predicate. A row group 
has two required INT32 columns:
   
   | Physical row | `predicate` | `payload` |
   | --- | ---: | ---: |
   | 0 | 0 | 0 |
   | 1 | 0 | 10 |
   | 2 | 100 | 20 |
   | 3 | 100 | 30 |
   | 4 | 0 | 40 |
   
   The columns have different, valid page boundaries: `predicate` has pages 
containing rows **[0,1]**, **[2,3]**, and **[4]**, while `payload` has pages 
containing rows **[0,1,2]** and **[3,4]**.
   
   Read this file with `ParquetReader<Group>`, 
`FilterCompat.get(eq(intColumn("predicate"), 0))`, and 
`useColumnIndexFilter(true)`. The column index excludes predicate page [2,3], 
leaving candidate row positions **0, 1, and 4**. The expected payloads are 
**[0, 10, 40]**, but the reader returns **[0, 10, 20]**: the last record 
combines the predicate value from row 4 with the payload from row 2. Reading 
the same file with column-index filtering disabled returns the correct result.
   
   ## Cause
   
   `SynchronizingColumnReader.isFullyConsumed()` currently checks only 
`!rowIndexes.hasNext()`. Consuming a position from this iterator means that the 
row has become the next target, not that its data has been read. In the 
example, the payload reader selects target row 4 while still on its first page. 
At the page transition, it exits instead of opening the second page and 
advancing to row 4.
   
   The completion check needs to account for the pending target as well as 
iterator exhaustion, while retaining the existing no-more-targets sentinel. 
This is a reader-state bug; it does not require a file-format change.
   
   ## Version and reproduction
   
   Reproduced against Apache Parquet `master` at 
`60175684378abff1ea001b6541ec38da42a2eff1` (`1.19.0-SNAPSHOT`), using OpenJDK 
17.0.19 on macOS arm64. A standalone check writes real Parquet files with the 
page boundaries above, then reads them through the normal predicate-filtering 
API. It uses no externally supplied row positions.
   
   The wrong result occurs with both V1 and V2 data pages and required INT32 
columns. Changing only the reader's completion check returns the correct result 
in both cases. Controls with aligned page boundaries, or with two consecutive 
rows in the final candidate range, already return the correct results on the 
old reader; not every filtered read triggers the problem.
   
   Focused column-reader regression tests also reproduce the failure with 
optional and repeated columns in both page versions: all four fail before the 
fix and pass afterward.
   
   ## Component(s)
   
   Core
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to