Aleksandr Efimov has posted comments on this change. ( 
http://gerrit.cloudera.org:8080/25002 )

Change subject: IMPALA-15059: Use HBO stats for Calcite filters
......................................................................


Patch Set 1:

(1 comment)

http://gerrit.cloudera.org:8080/#/c/25002/1/java/calcite-planner/src/main/java/org/apache/impala/calcite/schema/CalciteTable.java
File 
java/calcite-planner/src/main/java/org/apache/impala/calcite/schema/CalciteTable.java:

http://gerrit.cloudera.org:8080/#/c/25002/1/java/calcite-planner/src/main/java/org/apache/impala/calcite/schema/CalciteTable.java@108
PS1, Line 108:   private SimplifiedAnalyzer hboAnalyzer_;
> I will review this more closely (prolly in the coming week), but something
I tried a few ways to avoid the separate analyzer. Using the main analyzer 
added 13 slots and increased privilege requests from 2 to 14 for `select id ... 
where id = 1`. A non-materialized tuple avoids serialization and extra column 
requests, but still changes descriptors, aliases and expression counts.

Detached descriptors failed for complex pruning, e.g. `year + month = 2010`, 
because the pruner resolves slots through `analyzer.getDescTbl()`. An adapter 
with its own descriptors, expression counter and constant folder passed the 
focused checks, but depends more on Analyzer internals. Reusing the metadata 
TableRef for physical scans gave self-join operands one tuple ID and 
materialized 11 columns per scan instead of one. All 23 existing HBO tests 
still passed; additional invariant checks caught this.

A shared, isolated analyzer passed all 36 HBO tests in the join stack and 
matched lookup/pruning results for eight additional predicates. Focused SQL 
checks also returned identical results for 11 queries with HBO disabled, empty 
history and collected history. Fallback was disabled, and actual metadata hits 
were confirmed.

The planning benchmark showed no clear speed improvement. The shared prototype 
saved about 3% in metadata setup allocations, but allocated about 6% more 
during full three-table planning. It doesn't cache scan contexts across join 
alternatives, so this doesn't rule out a better query-scoped design.

My proposal is to keep a lazy, per-table metadata context for this patch, with 
final fields inside it. Laziness avoids unused setup; eager initialization 
passed too. Would that address your concern, or would you prefer query-scoped 
ownership?



--
To view, visit http://gerrit.cloudera.org:8080/25002
To unsubscribe, visit http://gerrit.cloudera.org:8080/settings

Gerrit-Project: Impala-ASF
Gerrit-Branch: master
Gerrit-MessageType: comment
Gerrit-Change-Id: I9270a5f94d42a4ad20d3337819b46467e83a3bf7
Gerrit-Change-Number: 25002
Gerrit-PatchSet: 1
Gerrit-Owner: Aleksandr Efimov <[email protected]>
Gerrit-Reviewer: Aleksandr Efimov <[email protected]>
Gerrit-Reviewer: Impala Public Jenkins <[email protected]>
Gerrit-Reviewer: Quanlong Huang <[email protected]>
Gerrit-Reviewer: Steve Carlin <[email protected]>
Gerrit-Comment-Date: Sat, 03 Oct 2026 22:34:37 +0000
Gerrit-HasComments: Yes

Reply via email to