Aleksandr Efimov has posted comments on this change. ( http://gerrit.cloudera.org:8080/25002 )
Change subject: IMPALA-15059: Use HBO stats for Calcite filters ...................................................................... Patch Set 1: (1 comment) http://gerrit.cloudera.org:8080/#/c/25002/1/java/calcite-planner/src/main/java/org/apache/impala/calcite/schema/CalciteTable.java File java/calcite-planner/src/main/java/org/apache/impala/calcite/schema/CalciteTable.java: http://gerrit.cloudera.org:8080/#/c/25002/1/java/calcite-planner/src/main/java/org/apache/impala/calcite/schema/CalciteTable.java@108 PS1, Line 108: private SimplifiedAnalyzer hboAnalyzer_; > I will review this more closely (prolly in the coming week), but something I tried a few ways to avoid the separate analyzer. Using the main analyzer added 13 slots and increased privilege requests from 2 to 14 for `select id ... where id = 1`. A non-materialized tuple avoids serialization and extra column requests, but still changes descriptors, aliases and expression counts. Detached descriptors failed for complex pruning, e.g. `year + month = 2010`, because the pruner resolves slots through `analyzer.getDescTbl()`. An adapter with its own descriptors, expression counter and constant folder passed the focused checks, but depends more on Analyzer internals. Reusing the metadata TableRef for physical scans gave self-join operands one tuple ID and materialized 11 columns per scan instead of one. All 23 existing HBO tests still passed; additional invariant checks caught this. A shared, isolated analyzer passed all 36 HBO tests in the join stack and matched lookup/pruning results for eight additional predicates. Focused SQL checks also returned identical results for 11 queries with HBO disabled, empty history and collected history. Fallback was disabled, and actual metadata hits were confirmed. The planning benchmark showed no clear speed improvement. The shared prototype saved about 3% in metadata setup allocations, but allocated about 6% more during full three-table planning. It doesn't cache scan contexts across join alternatives, so this doesn't rule out a better query-scoped design. My proposal is to keep a lazy, per-table metadata context for this patch, with final fields inside it. Laziness avoids unused setup; eager initialization passed too. Would that address your concern, or would you prefer query-scoped ownership? -- To view, visit http://gerrit.cloudera.org:8080/25002 To unsubscribe, visit http://gerrit.cloudera.org:8080/settings Gerrit-Project: Impala-ASF Gerrit-Branch: master Gerrit-MessageType: comment Gerrit-Change-Id: I9270a5f94d42a4ad20d3337819b46467e83a3bf7 Gerrit-Change-Number: 25002 Gerrit-PatchSet: 1 Gerrit-Owner: Aleksandr Efimov <[email protected]> Gerrit-Reviewer: Aleksandr Efimov <[email protected]> Gerrit-Reviewer: Impala Public Jenkins <[email protected]> Gerrit-Reviewer: Quanlong Huang <[email protected]> Gerrit-Reviewer: Steve Carlin <[email protected]> Gerrit-Comment-Date: Sat, 03 Oct 2026 22:34:37 +0000 Gerrit-HasComments: Yes
