HippoBaro opened a new issue, #11263:
URL: https://github.com/apache/arrow-rs/issues/11263
### Describe the bug
When mixing `write_batch` and `write_batch_with_statistics`, a data-page
flush can cause the writer to report a later batch's distinct count as the
distinct count for the entire column chunk.
Caller-supplied distinct counts describe an individual write. Once earlier
data makes the column-wide count unknown, a later write cannot safely restore
it using only its own batch statistics. However, the writer checks
`self.encoder.num_values() == 0` to decide whether any data has already been
written. That counter resets when a page is flushed, so an empty encoder is
incorrectly treated as an empty column chunk.
For example, writing `[1, 2, 3, 4]` without a precomputed distinct count,
followed by `[5, 6, 7]` with a supplied count of `3`, can produce column
metadata claiming that the whole chunk has only **3 distinct values**. The
actual values are unchanged and have **7 distinct values**. Without a page
boundary between those writes, the distinct count correctly remains absent.
### To Reproduce
```rust
use std::{fs::File, sync::Arc};
use parquet::{
data_type::Int32Type,
file::{
properties::WriterProperties,
reader::{FileReader, SerializedFileReader},
writer::SerializedFileWriter,
},
schema::parser::parse_message_type,
};
fn main() -> Result<(), Box<dyn std::error::Error>> {
for page_limit in [100, 4] {
let path = format!("distinct-count-{page_limit}.parquet");
let schema = Arc::new(parse_message_type("message schema { REQUIRED
INT32 value; }")?);
let props = Arc::new(
WriterProperties::builder()
.set_write_page_header_statistics(true)
.set_data_page_row_count_limit(page_limit)
.build(),
);
let mut writer = SerializedFileWriter::new(File::create(&path)?,
schema, props)?;
let mut row_group = writer.next_row_group()?;
let mut column = row_group.next_column()?.unwrap();
let typed = column.typed::<Int32Type>();
// No precomputed distinct count for the first write.
typed.write_batch(&[1, 2, 3, 4], None, None)?;
// This distinct count describes only the second write.
typed.write_batch_with_statistics(&[5, 6, 7], None, None, Some(&5),
Some(&7), Some(3))?;
column.close()?;
row_group.close()?;
writer.close()?;
let reader = SerializedFileReader::new(File::open(&path)?)?;
let chunk = reader.metadata().row_group(0).column(0);
println!(
"page limit={page_limit}, values={}, distinct_count={:?}",
chunk.num_values(),
chunk.statistics().unwrap().distinct_count_opt(),
);
}
Ok(())
}
```
Actual output:
```text
page limit=100, values=7, distinct_count=None
page limit=4, values=7, distinct_count=Some(3)
```
### Expected behavior
The column-chunk distinct count should remain absent in both cases,
regardless of where page boundaries fall:
```text
page limit=100, values=7, distinct_count=None
page limit=4, values=7, distinct_count=None
```
The writer should only accept the supplied batch distinct count when there
is no prior data in the column chunk. A page flush must not erase that history
or allow a later batch to restore an unknown count. This does not require the
writer to compute the actual count of `7`; it should omit a count it cannot
establish safely. Summing batch distinct counts is not generally valid either,
since values can overlap across writes.
### Additional context
_No response_
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]