HippoBaro opened a new issue, #11263:
URL: https://github.com/apache/arrow-rs/issues/11263

   ### Describe the bug
   
   When mixing `write_batch` and `write_batch_with_statistics`, a data-page 
flush can cause the writer to report a later batch's distinct count as the 
distinct count for the entire column chunk.
   
   Caller-supplied distinct counts describe an individual write. Once earlier 
data makes the column-wide count unknown, a later write cannot safely restore 
it using only its own batch statistics. However, the writer checks 
`self.encoder.num_values() == 0` to decide whether any data has already been 
written. That counter resets when a page is flushed, so an empty encoder is 
incorrectly treated as an empty column chunk.
   
   For example, writing `[1, 2, 3, 4]` without a precomputed distinct count, 
followed by `[5, 6, 7]` with a supplied count of `3`, can produce column 
metadata claiming that the whole chunk has only **3 distinct values**. The 
actual values are unchanged and have **7 distinct values**. Without a page 
boundary between those writes, the distinct count correctly remains absent.
   
   
   ### To Reproduce
   
   ```rust
   use std::{fs::File, sync::Arc};
   
   use parquet::{
       data_type::Int32Type,
       file::{
           properties::WriterProperties,
           reader::{FileReader, SerializedFileReader},
           writer::SerializedFileWriter,
       },
       schema::parser::parse_message_type,
   };
   
   fn main() -> Result<(), Box<dyn std::error::Error>> {
       for page_limit in [100, 4] {
           let path = format!("distinct-count-{page_limit}.parquet");
           let schema = Arc::new(parse_message_type("message schema { REQUIRED 
INT32 value; }")?);
           let props = Arc::new(
               WriterProperties::builder()
                   .set_write_page_header_statistics(true)
                   .set_data_page_row_count_limit(page_limit)
                   .build(),
           );
           let mut writer = SerializedFileWriter::new(File::create(&path)?, 
schema, props)?;
           let mut row_group = writer.next_row_group()?;
           let mut column = row_group.next_column()?.unwrap();
           let typed = column.typed::<Int32Type>();
   
           // No precomputed distinct count for the first write.
           typed.write_batch(&[1, 2, 3, 4], None, None)?;
           // This distinct count describes only the second write.
           typed.write_batch_with_statistics(&[5, 6, 7], None, None, Some(&5), 
Some(&7), Some(3))?;
   
           column.close()?;
           row_group.close()?;
           writer.close()?;
   
           let reader = SerializedFileReader::new(File::open(&path)?)?;
           let chunk = reader.metadata().row_group(0).column(0);
           println!(
               "page limit={page_limit}, values={}, distinct_count={:?}",
               chunk.num_values(),
               chunk.statistics().unwrap().distinct_count_opt(),
           );
       }
       Ok(())
   }
   ```
   
   Actual output:
   
   ```text
   page limit=100, values=7, distinct_count=None
   page limit=4, values=7, distinct_count=Some(3)
   ```
   
   ### Expected behavior
   
   The column-chunk distinct count should remain absent in both cases, 
regardless of where page boundaries fall:
   
   ```text
   page limit=100, values=7, distinct_count=None
   page limit=4, values=7, distinct_count=None
   ```
   
   The writer should only accept the supplied batch distinct count when there 
is no prior data in the column chunk. A page flush must not erase that history 
or allow a later batch to restore an unknown count. This does not require the 
writer to compute the actual count of `7`; it should omit a count it cannot 
establish safely. Summing batch distinct counts is not generally valid either, 
since values can overlap across writes.
   
   
   ### Additional context
   
   _No response_


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to