Hello Quanlong Huang, Michael Smith, Impala Public Jenkins,

I'd like you to reexamine a change. Please visit

    http://gerrit.cloudera.org:8080/24913

to look at the new patch set (#4).

Change subject: IMPALA-15394: Add missing scan properties to the HBO scan key
......................................................................

IMPALA-15394: Add missing scan properties to the HBO scan key

HdfsScanNode.generateHboKeyString() keys a scan by its table and
conjuncts. Some scans of the same table and conjuncts return different
rows, and their input stats match too, so HBO gives one the cardinality
of the other:
- An optimized count(*) scan returns one row per file or row group.
  After "SELECT count(*) FROM t" on a Parquet table with 15 files and
  13.66M rows, "SELECT * FROM t" is planned with cardinality=15 (from
  HBO), and every estimate above the scan follows.
- A partition key scan, e.g. "SELECT DISTINCT year FROM t", returns one
  row per scan range.
- A sampled scan (TABLESAMPLE) reads only part of the files, while its
  input rows count every partition with a sampled file (or all files of
  an Iceberg table). Without numRows, runs match by catalog version.
- Conjuncts on the items of a collection, together with the
  IsNotEmptyPredicate of the join, filter the rows of the scan. On
  functional_parquet.complextypestbl the scan returned 2 rows for
  "a.item > 1" and 0 for "a.item > 100" under the same key.

This patch adds <COUNT_STAR>, <PARTITION_KEY_SCAN> and <SAMPLE:percent>
markers and the canonicalized collection conjuncts to the scan key. The
collection is named by its canonical path, so the key doesn't depend on
the aliases. The seed of a sample is left out: it picks other files but
reads the same share of the table.

Testing:
- Added HboKeyStringTest.testCountStarScanNodeKey,
  testPartitionKeyAndSampledScanNodeKeys and
  testCollectionConjunctsInScanNodeKey. Removing each marker or the
  collection conjuncts, or naming the collection by its alias, fails
  the test that covers it.
- Added e2e cases to test_hbo.py: a regular scan after a count(*) scan,
  and item conjuncts of a joined collection. Without this patch both
  get the cardinality of the other query from HBO.

Change-Id: I03183aa998f90bf834dba4ce9ad2f2e57fef1140
Assisted-by: claude-opus-5 (Claude Code)
---
M fe/src/main/java/org/apache/impala/planner/HdfsScanNode.java
M fe/src/test/java/org/apache/impala/planner/HboKeyStringTest.java
M testdata/workloads/functional-query/queries/QueryTest/hbo-collection-scan.test
A testdata/workloads/functional-query/queries/QueryTest/hbo-count-star-scan.test
M tests/query_test/test_hbo.py
5 files changed, 199 insertions(+), 0 deletions(-)


  git pull ssh://gerrit.cloudera.org:29418/Impala-ASF refs/changes/13/24913/4
-- 
To view, visit http://gerrit.cloudera.org:8080/24913
To unsubscribe, visit http://gerrit.cloudera.org:8080/settings

Gerrit-Project: Impala-ASF
Gerrit-Branch: master
Gerrit-MessageType: newpatchset
Gerrit-Change-Id: I03183aa998f90bf834dba4ce9ad2f2e57fef1140
Gerrit-Change-Number: 24913
Gerrit-PatchSet: 4
Gerrit-Owner: Aleksandr Efimov <[email protected]>
Gerrit-Reviewer: Aleksandr Efimov <[email protected]>
Gerrit-Reviewer: Impala Public Jenkins <[email protected]>
Gerrit-Reviewer: Michael Smith <[email protected]>
Gerrit-Reviewer: Quanlong Huang <[email protected]>

Reply via email to