xudong963 opened a new pull request, #25854:
URL: https://github.com/apache/datafusion/pull/25854

   ## Which issue does this PR close?
   
   No linked issue. This is a follow-up to the fully matched Parquet row-group 
work in PR #23696.
   
   ## Rationale for this change
   
   When row-group statistics prove that every row satisfies a scan predicate, a 
Bloom filter cannot prune that group. Reading its Bloom filters still adds 
object-store I/O, which can be especially costly for remote files.
   
   ## What changes are included in this PR?
   
   - Skip Bloom filter reads and predicate evaluation for fully matched row 
groups.
   - Avoid creating a Bloom reader when every surviving row group is fully 
matched.
   - Keep the Bloom pruning matched metric accounting for skipped groups.
   
   ## What is the testing strategy for this PR?
   
   The new `fully_matched_row_groups_skip_bloom_filter_reads` test verifies 
that a partially matched group still reads Bloom filters, a fully matched group 
reduces `bytes_scanned`, and an all-fully-matched file reads zero Bloom bytes 
during open. It also checks that the returned rows are unchanged.
   
   Validation:
   
   - `cargo fmt --all`
   - `cargo clippy --all-targets --all-features -- -D warnings`
   - `./dev/rust_lint.sh`
   - `RUST_BACKTRACE=1 cargo test --profile ci --exclude datafusion-examples 
--exclude datafusion-benchmarks --exclude datafusion-cli --workspace --lib 
--tests --bins --features 
avro,json,backtrace,extended_tests,recursive_protection,parquet_encryption`
   
   The test establishes reduced Bloom I/O. Remote object-store latency has not 
been benchmarked; the existing Parquet row-filter benchmark does not generate 
Bloom filters or use a Bloom-eligible equality predicate.
   
   ## Are there any user-facing changes?
   
   No API or query-result changes. Scans avoid unnecessary Bloom filter reads 
for fully matched row groups.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to