This is an automated email from the ASF dual-hosted git repository.

tuhaihe pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/cloudberry.git


The following commit(s) were added to refs/heads/main by this push:
     new 8fdd5caec71 [ORCA] Support hash partitioning in ORCA
8fdd5caec71 is described below

commit 8fdd5caec7105e0f522ad9931d137b8211eb7eaa
Author: Zhang Wenchao <[email protected]>
AuthorDate: Mon Jun 22 18:24:29 2026 +0800

    [ORCA] Support hash partitioning in ORCA
    
    GPORCA previously rejected any query touching a hash-partitioned table
    and fell back to the Postgres planner. Enable GPORCA to plan such
    queries, with static partition pruning for equality predicates and
    dynamic (join-driven) partition elimination.
    
    How:
    - CTranslatorRelcacheToDXL: stop raising "hash partitioning" for a
      single-column, single-level hash partition key and let the strategy
      flow into GPORCA. Composite keys, partitioning by expression and
      multi-level partitioning still fall back to the planner.
    - IMDRelation: add ErelpartitionHash ('h') to the partition-type enum.
    - CExpressionPreprocessor: a hash leaf's qual is a
      satisfies_hash_partition() call with no btree-interval form, so
      PcnstrFromChildPartition now returns NULL instead of asserting. Add
      FHashPartitionPruned(): substitute the query's equality constants for
      the partition-key columns into that call and evaluate it with the
      constant-expression evaluator; a leaf whose call folds to false cannot
      hold a matching row and is pruned. This reuses PostgreSQL's exact
      hashing (seed, per-column proc, combine), so selection is identical.
    - CConstExprEvaluatorDXL: allow folding any column-free immutable
      expression (e.g. satisfies_hash_partition over constants), not only
      (const cmp const). Volatile/stable and column-referencing expressions
      are still rejected.
    
    Dynamic partition elimination needs no hash-specific code: once hash
    tables become CLogicalDynamicGet the existing CPhysicalPartitionSelector
    path is partition-type agnostic, and CPartPruneStepsBuilder resolves the
    strategy from the partition's own hash opfamily.
    
    ORCA vs Postgres planner behavior:
    - Partition selection and results are identical. For the supported
      single-column, single-level hash key, an equality predicate prunes to
      exactly the same surviving partition(s) as the Postgres planner and
      returns the same rows row-for-row. When a single partition survives,
      both narrow dispatch to that one segment (Gather Motion 1:1).
    - Unchanged: composite keys, partitioning by expression and multi-level
      partitioning still fall back to the Postgres planner, and range/list
      pruning is unaffected.
    
    Authored-by: Zhang Wenchao <[email protected]>
---
 .../regress/expected/direct_dispatch_optimizer.out |  14 +-
 .../expected/orca_static_pruning_optimizer.out     |  15 +-
 .../regress/expected/partition_prune_optimizer.out |  28 +--
 .../gpopt/translate/CTranslatorRelcacheToDXL.cpp   |  13 +-
 .../gpopt/operators/CExpressionPreprocessor.h      |  18 ++
 .../libgpopt/src/eval/CConstExprEvaluatorDXL.cpp   |  12 +-
 .../src/operators/CExpressionPreprocessor.cpp      | 274 +++++++++++++++++++--
 .../include/naucrates/md/IMDRelation.h             |   3 +-
 .../regress/expected/direct_dispatch_optimizer.out |  14 +-
 .../expected/orca_static_pruning_optimizer.out     |  15 +-
 .../regress/expected/partition_prune_append.out    |  26 +-
 .../regress/expected/partition_prune_optimizer.out |  28 +--
 .../regress/expected/qp_indexscan_optimizer.out    |  15 +-
 13 files changed, 362 insertions(+), 113 deletions(-)

diff --git 
a/contrib/pax_storage/src/test/regress/expected/direct_dispatch_optimizer.out 
b/contrib/pax_storage/src/test/regress/expected/direct_dispatch_optimizer.out
index 8efc7ccfc0f..d6fd64b18f4 100644
--- 
a/contrib/pax_storage/src/test/regress/expected/direct_dispatch_optimizer.out
+++ 
b/contrib/pax_storage/src/test/regress/expected/direct_dispatch_optimizer.out
@@ -1665,15 +1665,11 @@ explain select * from t_14887 where a = 'a'::text;
                                     QUERY PLAN                                 
   
 
----------------------------------------------------------------------------------
  Gather Motion 3:1  (slice1; segments: 3)  (cost=0.00..764.38 rows=158 
width=32)
-   ->  Append  (cost=0.00..762.26 rows=53 width=32)
-         ->  Seq Scan on t0_14887 t_14887_1  (cost=0.00..254.00 rows=18 
width=32)
-               Filter: ((a)::text = 'a'::text)
-         ->  Seq Scan on t1_14887 t_14887_2  (cost=0.00..254.00 rows=18 
width=32)
-               Filter: ((a)::text = 'a'::text)
-         ->  Seq Scan on t2_14887 t_14887_3  (cost=0.00..254.00 rows=18 
width=32)
-               Filter: ((a)::text = 'a'::text)
- Optimizer: Postgres query optimizer
-(9 rows)
+   ->  Dynamic Seq Scan on t_14887
+         Number of partitions to scan: 3 (out of 3)
+         Filter: ((a)::text = 'a'::text)
+ Optimizer: Pivotal Optimizer (GPORCA)
+(4 rows)
 
 begin;
 drop table if exists direct_test;
diff --git 
a/contrib/pax_storage/src/test/regress/expected/orca_static_pruning_optimizer.out
 
b/contrib/pax_storage/src/test/regress/expected/orca_static_pruning_optimizer.out
index 15ffbe07d84..4dd130ab339 100644
--- 
a/contrib/pax_storage/src/test/regress/expected/orca_static_pruning_optimizer.out
+++ 
b/contrib/pax_storage/src/test/regress/expected/orca_static_pruning_optimizer.out
@@ -160,21 +160,18 @@ WHERE b = 42
 $query$ AS qry \gset
 EXPLAIN (COSTS OFF, VERBOSE)
 :qry ;
-INFO:  GPORCA failed to produce a plan, falling back to Postgres-based planner
-DETAIL:  Falling back to Postgres-based planner because GPORCA does not 
support the following feature: hash partitioning
                   QUERY PLAN                  
 ----------------------------------------------
  Gather Motion 3:1  (slice1; segments: 3)
-   Output: hp.a, hp.b
-   ->  Seq Scan on orca_static_pruning.hp0 hp
-         Output: hp.a, hp.b
+   Output: a, b
+   ->  Dynamic Seq Scan on orca_static_pruning.hp
+         Output: a, b
+         Number of partitions to scan: 1 (out of 2)
          Filter: (hp.b = 42)
- Optimizer: Postgres query optimizer
-(7 rows)
+ Optimizer: Pivotal Optimizer (GPORCA)
+(8 rows)
 
 :qry ;
-INFO:  GPORCA failed to produce a plan, falling back to Postgres-based planner
-DETAIL:  Falling back to Postgres-based planner because GPORCA does not 
support the following feature: hash partitioning
  a | b  
 ---+----
  0 | 42
diff --git 
a/contrib/pax_storage/src/test/regress/expected/partition_prune_optimizer.out 
b/contrib/pax_storage/src/test/regress/expected/partition_prune_optimizer.out
index 58a9d806d73..f5303b87634 100644
--- 
a/contrib/pax_storage/src/test/regress/expected/partition_prune_optimizer.out
+++ 
b/contrib/pax_storage/src/test/regress/expected/partition_prune_optimizer.out
@@ -4313,31 +4313,31 @@ explain (costs off) select * from pph_arrpart where a = 
'{1}';
                  QUERY PLAN                 
 --------------------------------------------
  Gather Motion 1:1  (slice1; segments: 1)
-   ->  Seq Scan on pph_arrpart2 pph_arrpart
+   ->  Dynamic Seq Scan on pph_arrpart
+         Number of partitions to scan: 1 (out of 2)
          Filter: (a = '{1}'::integer[])
- Optimizer: Postgres query optimizer
-(4 rows)
+ Optimizer: Pivotal Optimizer (GPORCA)
+(5 rows)
 
 explain (costs off) select * from pph_arrpart where a = '{1, 2}';
                  QUERY PLAN                 
 --------------------------------------------
  Gather Motion 1:1  (slice1; segments: 1)
-   ->  Seq Scan on pph_arrpart1 pph_arrpart
+   ->  Dynamic Seq Scan on pph_arrpart
+         Number of partitions to scan: 1 (out of 2)
          Filter: (a = '{1,2}'::integer[])
- Optimizer: Postgres query optimizer
-(4 rows)
+ Optimizer: Pivotal Optimizer (GPORCA)
+(5 rows)
 
 explain (costs off) select * from pph_arrpart where a in ('{4, 5}', '{1}');
                                  QUERY PLAN                                 
 ----------------------------------------------------------------------------
- Gather Motion 2:1  (slice1; segments: 2)
-   ->  Append
-         ->  Seq Scan on pph_arrpart1 pph_arrpart_1
-               Filter: ((a = '{4,5}'::integer[]) OR (a = '{1}'::integer[]))
-         ->  Seq Scan on pph_arrpart2 pph_arrpart_2
-               Filter: ((a = '{4,5}'::integer[]) OR (a = '{1}'::integer[]))
- Optimizer: Postgres query optimizer
-(7 rows)
+ Gather Motion 3:1  (slice1; segments: 3)
+   ->  Dynamic Seq Scan on pph_arrpart
+         Number of partitions to scan: 2 (out of 2)
+         Filter: ((a = '{4,5}'::integer[]) OR (a = '{1}'::integer[]))
+ Optimizer: Pivotal Optimizer (GPORCA)
+(5 rows)
 
 drop table pph_arrpart;
 -- enum type list partition key
diff --git a/src/backend/gpopt/translate/CTranslatorRelcacheToDXL.cpp 
b/src/backend/gpopt/translate/CTranslatorRelcacheToDXL.cpp
index aa23d3932a2..119dc18f4df 100644
--- a/src/backend/gpopt/translate/CTranslatorRelcacheToDXL.cpp
+++ b/src/backend/gpopt/translate/CTranslatorRelcacheToDXL.cpp
@@ -2692,11 +2692,14 @@ 
CTranslatorRelcacheToDXL::RetrievePartKeysAndTypes(CMemoryPool *mp,
                                   GPOS_WSZ_LIT("partitioning by expression"));
        }
 
-       if (PARTITION_STRATEGY_HASH == part_type)
-       {
-               GPOS_RAISE(gpdxl::ExmaMD, gpdxl::ExmiMDObjUnsupported,
-                                  GPOS_WSZ_LIT("hash partitioning"));
-       }
+       // Hash-partitioned tables are supported for scanning only. ORCA cannot
+       // express a hash partition's membership rule (the 
satisfies_hash_partition
+       // call, i.e. hash(key) mod modulus = remainder) as a btree interval
+       // constraint, so it does no static partition pruning for hash 
partitions
+       // and simply scans every leaf partition. Range and list partitions keep
+       // full pruning support. Letting the strategy flow through here 
(instead of
+       // raising ExmiMDObjUnsupported) keeps such queries inside ORCA rather 
than
+       // falling back to the Postgres planner.
 
        (*part_keys)->Append(GPOS_NEW(mp) ULONG(attno - 1));
        (*part_types)->Append(GPOS_NEW(mp) CHAR(part_type));
diff --git 
a/src/backend/gporca/libgpopt/include/gpopt/operators/CExpressionPreprocessor.h 
b/src/backend/gporca/libgpopt/include/gpopt/operators/CExpressionPreprocessor.h
index 978d2575e42..58965fb2e9b 100644
--- 
a/src/backend/gporca/libgpopt/include/gpopt/operators/CExpressionPreprocessor.h
+++ 
b/src/backend/gporca/libgpopt/include/gpopt/operators/CExpressionPreprocessor.h
@@ -215,6 +215,24 @@ private:
                                                                                
                 CColRefArray *pdrgpcrOutput,
                                                                                
                 ColRefToUlongMap *col_mapping);
 
+       // translate a child partition's part-constraint DXL into a CExpression 
over
+       // the root table's colrefs (NULL if the child has no stored constraint)
+       static CExpression *PexprPartConstraintFromChild(
+               CMemoryPool *mp, const IMDRelation *partrel, CColRefArray 
*pdrgpcrOutput,
+               ColRefToUlongMap *root_col_mapping);
+
+       // find a "colref = const" equality among the given conjuncts and 
return its
+       // const subexpression (NULL if none)
+       static CExpression *PexprColumnEqualityConst(
+               CExpressionArray *pdrgpexprConjuncts, const CColRef *colref);
+
+       // static hash-partition pruning: true if the leaf provably cannot 
contain
+       // any row matching the query's equality predicates on the partition key
+       static BOOL FHashPartitionPruned(CMemoryPool *mp, const IMDRelation 
*partrel,
+                                                                        
CColRefArray *pdrgpcrOutput,
+                                                                        
ColRefToUlongMap *root_col_mapping,
+                                                                        
CExpressionArray *pdrgpexprConjuncts);
+
        // swap logical select over logical project
        static CExpression *PexprTransposeSelectAndProject(CMemoryPool *mp,
                                                                                
                           CExpression *pexpr);
diff --git a/src/backend/gporca/libgpopt/src/eval/CConstExprEvaluatorDXL.cpp 
b/src/backend/gporca/libgpopt/src/eval/CConstExprEvaluatorDXL.cpp
index 80ac01b8482..47d90f0e101 100644
--- a/src/backend/gporca/libgpopt/src/eval/CConstExprEvaluatorDXL.cpp
+++ b/src/backend/gporca/libgpopt/src/eval/CConstExprEvaluatorDXL.cpp
@@ -24,6 +24,8 @@
 #include "gpopt/operators/CExpression.h"
 #include "gpopt/operators/CPredicateUtils.h"
 
+#include "naucrates/md/IMDFunction.h"
+
 using namespace gpdxl;
 using namespace gpmd;
 using namespace gpopt;
@@ -72,7 +74,15 @@ CConstExprEvaluatorDXL::PexprEval(CExpression *pexpr)
 {
        GPOS_ASSERT(nullptr != pexpr);
 
-       if (!CPredicateUtils::FCompareConstToConstIgnoreCast(pexpr))
+       // We can evaluate (a) the (const cmp const) expressions used by the
+       // comparator, and (b) any column-free IMMUTABLE expression (e.g.
+       // satisfies_hash_partition() over constants, used for hash partition
+       // pruning). We must not fold expressions that reference columns, nor
+       // volatile/stable ones, since their value is not fixed at plan time.
+       if (!(CPredicateUtils::FCompareConstToConstIgnoreCast(pexpr) ||
+                 (0 == pexpr->DeriveUsedColumns()->Size() &&
+                  IMDFunction::EfsImmutable ==
+                          pexpr->DeriveScalarFunctionProperties()->Efs())))
        {
                GPOS_RAISE(gpopt::ExmaGPOPT, 
gpopt::ExmiEvalUnsupportedScalarExpr);
        }
diff --git 
a/src/backend/gporca/libgpopt/src/operators/CExpressionPreprocessor.cpp 
b/src/backend/gporca/libgpopt/src/operators/CExpressionPreprocessor.cpp
index e1c46c9fc8e..6755cc0714e 100644
--- a/src/backend/gporca/libgpopt/src/operators/CExpressionPreprocessor.cpp
+++ b/src/backend/gporca/libgpopt/src/operators/CExpressionPreprocessor.cpp
@@ -22,6 +22,7 @@
 #include "gpopt/base/CConstraintInterval.h"
 #include "gpopt/base/COptCtxt.h"
 #include "gpopt/base/CUtils.h"
+#include "gpopt/eval/IConstExprEvaluator.h"
 #include "gpopt/exception.h"
 #include "gpopt/mdcache/CMDAccessor.h"
 #include "gpopt/operators/CDedupSupersetPreprocessor.h"
@@ -49,6 +50,9 @@
 #include "gpopt/operators/COrderedAggPreprocessor.h"
 #include "gpopt/operators/CPredicateUtils.h"
 #include "gpopt/operators/CScalarCmp.h"
+#include "gpopt/operators/CScalarConst.h"
+#include "gpopt/operators/CScalarFunc.h"
+#include "gpopt/operators/CScalarIdent.h"
 #include "gpopt/operators/CScalarNAryJoinPredList.h"
 #include "gpopt/operators/CScalarProjectElement.h"
 #include "gpopt/operators/CScalarProjectList.h"
@@ -59,6 +63,7 @@
 #include "gpopt/optimizer/COptimizerConfig.h"
 #include "gpopt/translate/CTranslatorDXLToExpr.h"
 #include "gpopt/xforms/CXform.h"
+#include "naucrates/base/IDatumBool.h"
 #include "naucrates/md/IMDScalarOp.h"
 #include "naucrates/md/IMDType.h"
 #include "naucrates/statistics/CStatistics.h"
@@ -2848,13 +2853,29 @@ CExpressionPreprocessor::PrunePartitions(CMemoryPool 
*mp, CExpression *expr)
                        CLogicalDynamicGet::PopConvert((*expr)[0]->Pop());
 
                CColRefSetArray *pdrgpcrsChild = nullptr;
-               // As of now, partition's default opfamily is btree
-               // ORCA doesn't support hash partition yet
+               // pred_cnstr captures predicates on the partition key as a 
btree interval
+               // constraint, used to statically prune range/list partitions. 
Hash
+               // partitions have no btree-interval form and are pruned 
separately below
+               // via FHashPartitionPruned().
                CConstraint *pred_cnstr = CConstraint::PcnstrFromScalarExpr(
                        mp, filter_pred, &pdrgpcrsChild, false /* 
infer_nulls_as*/,
                        IMDIndex::EmdindBtree);
                CRefCount::SafeRelease(pdrgpcrsChild);
 
+               // For hash-partitioned tables ORCA cannot build btree interval
+               // constraints, so static pruning is done separately by 
evaluating each
+               // leaf's satisfies_hash_partition() qual against the query's 
equality
+               // predicates (see FHashPartitionPruned).
+               const IMDRelation *root_rel =
+                       mda->RetrieveRel(dyn_get->Ptabdesc()->MDId());
+               BOOL fHashPartitioned =
+                       root_rel->IsPartitioned() &&
+                       IMDRelation::ErelpartitionHash == 
root_rel->PartTypeAtLevel(0);
+               CExpressionArray *pdrgpexprConjuncts =
+                       fHashPartitioned
+                               ? CPredicateUtils::PdrgpexprConjuncts(mp, 
filter_pred)
+                               : nullptr;
+
                IMdIdArray *selected_partition_mdids = GPOS_NEW(mp) 
IMdIdArray(mp);
                CConstraintArray *selected_partition_cnstrs =
                        GPOS_NEW(mp) CConstraintArray(mp);
@@ -2866,9 +2887,26 @@ CExpressionPreprocessor::PrunePartitions(CMemoryPool 
*mp, CExpression *expr)
                        IMDId *part_mdid = (*all_partition_mdids)[ul];
                        const IMDRelation *partrel = 
mda->RetrieveRel(part_mdid);
 
-                       CConstraint *rel_cnstr = PcnstrFromChildPartition(
-                               partrel, dyn_get->PdrgpcrOutput(),
-                               (*dyn_get->GetRootColMappingPerPart())[ul]);
+                       if (fHashPartitioned &&
+                               FHashPartitionPruned(
+                                       mp, partrel, dyn_get->PdrgpcrOutput(),
+                                       
(*dyn_get->GetRootColMappingPerPart())[ul],
+                                       pdrgpexprConjuncts))
+                       {
+                               // this hash leaf cannot match the equality 
predicates: prune it
+                               continue;
+                       }
+
+                       // For hash partitions PcnstrFromChildPartition always 
returns a null
+                       // btree constraint (the qual is 
satisfies_hash_partition()), so skip
+                       // the wasted DXL translation and keep rel_cnstr null.
+                       CConstraint *rel_cnstr = nullptr;
+                       if (!fHashPartitioned)
+                       {
+                               rel_cnstr = PcnstrFromChildPartition(
+                                       partrel, dyn_get->PdrgpcrOutput(),
+                                       
(*dyn_get->GetRootColMappingPerPart())[ul]);
+                       }
 
                        CConstraint *pcnstr = nullptr;
                        {
@@ -2895,9 +2933,12 @@ CExpressionPreprocessor::PrunePartitions(CMemoryPool 
*mp, CExpression *expr)
                                
foreign_server_mdids->Append(foreign_server_mdid);
                                part_mdid->AddRef();
                                selected_partition_mdids->Append(part_mdid);
-                               rel_cnstr = PcnstrFromChildPartition(
-                                       partrel, dyn_get->PdrgpcrOutput(),
-                                       
(*dyn_get->GetRootColMappingPerPart())[ul]);
+                               if (!fHashPartitioned)
+                               {
+                                       rel_cnstr = PcnstrFromChildPartition(
+                                               partrel, 
dyn_get->PdrgpcrOutput(),
+                                               
(*dyn_get->GetRootColMappingPerPart())[ul]);
+                               }
                                if (rel_cnstr)
                                {
                                        
selected_partition_cnstrs->Append(rel_cnstr);
@@ -2906,6 +2947,7 @@ CExpressionPreprocessor::PrunePartitions(CMemoryPool *mp, 
CExpression *expr)
                        CRefCount::SafeRelease(pcnstr);
                }
                CRefCount::SafeRelease(pred_cnstr);
+               CRefCount::SafeRelease(pdrgpexprConjuncts);
 
                if (selected_partition_mdids->Size() == 0)
                {
@@ -2958,21 +3000,18 @@ CExpressionPreprocessor::PrunePartitions(CMemoryPool 
*mp, CExpression *expr)
        return GPOS_NEW(mp) CExpression(mp, pop, children);
 }
 
-// Translate the part constraint of a child partition into an ORCA expr using
-// corresponding colrefs of the root table, instead of those from the child
-// partition.
-CConstraint *
-CExpressionPreprocessor::PcnstrFromChildPartition(
-       const IMDRelation *partrel, CColRefArray *pdrgpcrOutput,
+// Translate a child partition's stored part-constraint DXL into an ORCA
+// CExpression using the corresponding colrefs of the root table, instead of
+// those from the child partition. Returns NULL if the child has no stored
+// constraint (e.g. a default partition).
+CExpression *
+CExpressionPreprocessor::PexprPartConstraintFromChild(
+       CMemoryPool *mp, const IMDRelation *partrel, CColRefArray 
*pdrgpcrOutput,
        ColRefToUlongMap *root_col_mapping)
 {
        CMDAccessor *md_accessor = COptCtxt::PoctxtFromTLS()->Pmda();
-       CMemoryPool *mp = COptCtxt::PoctxtFromTLS()->Pmp();
-
-       CExpression *part_constraint_expr = nullptr;
 
        CDXLNode *dxlnode = partrel->MDPartConstraint();
-
        if (nullptr == dxlnode)
        {
                return nullptr;
@@ -2995,25 +3034,216 @@ CExpressionPreprocessor::PcnstrFromChildPartition(
        }
 
        CTranslatorDXLToExpr dxltr(mp, md_accessor);
-       part_constraint_expr =
+       CExpression *part_constraint_expr =
                dxltr.PexprTranslateScalar(dxlnode, pdrgpcrOutput, 
mapped_colids);
        mapped_colids->Release();
 
+       return part_constraint_expr;
+}
+
+// Build a CConstraint from a child partition's part constraint, used for 
static
+// pruning of range/list partitions.
+CConstraint *
+CExpressionPreprocessor::PcnstrFromChildPartition(
+       const IMDRelation *partrel, CColRefArray *pdrgpcrOutput,
+       ColRefToUlongMap *root_col_mapping)
+{
+       CMemoryPool *mp = COptCtxt::PoctxtFromTLS()->Pmp();
+
+       CExpression *part_constraint_expr = PexprPartConstraintFromChild(
+               mp, partrel, pdrgpcrOutput, root_col_mapping);
+       if (nullptr == part_constraint_expr)
+       {
+               return nullptr;
+       }
+
        GPOS_ASSERT(CUtils::FPredicate(part_constraint_expr));
 
        CColRefSetArray *pdrgpcrsChild = nullptr;
        CConstraint *cnstr;
-       // As of now, partition's default opfamily is btree
-       // ORCA doesn't support hash partition yet
+       // Partition constraints are extracted using the btree opfamily, which
+       // covers range and list partitions (their bounds are btree comparison
+       // predicates). A hash partition's constraint is a 
satisfies_hash_partition()
+       // function call that has no btree-interval representation, so
+       // PcnstrFromScalarExpr returns null for it; such partitions are pruned 
(when
+       // possible) by FHashPartitionPruned() instead.
        cnstr = CConstraint::PcnstrFromScalarExpr(
                mp, part_constraint_expr, &pdrgpcrsChild, true /* 
infer_nulls_as */,
                IMDIndex::EmdindBtree);
        CRefCount::SafeRelease(part_constraint_expr);
        CRefCount::SafeRelease(pdrgpcrsChild);
-       GPOS_ASSERT(cnstr);
        return cnstr;
 }
 
+// Return the const subexpression of a "colref = const" equality found among 
the
+// given conjuncts (NULL if none). The returned expression is owned by the
+// conjunct array; the caller must AddRef it before reusing it.
+CExpression *
+CExpressionPreprocessor::PexprColumnEqualityConst(
+       CExpressionArray *pdrgpexprConjuncts, const CColRef *colref)
+{
+       const ULONG size = pdrgpexprConjuncts->Size();
+       for (ULONG ul = 0; ul < size; ++ul)
+       {
+               CExpression *pexpr = (*pdrgpexprConjuncts)[ul];
+               if 
(!CPredicateUtils::FPlainEqualityIdentConstWithoutCast(pexpr))
+               {
+                       continue;
+               }
+
+               CExpression *pexprLeft = (*pexpr)[0];
+               CExpression *pexprRight = (*pexpr)[1];
+               CExpression *pexprIdent = nullptr;
+               CExpression *pexprConst = nullptr;
+               if (COperator::EopScalarIdent == pexprLeft->Pop()->Eopid() &&
+                       COperator::EopScalarConst == pexprRight->Pop()->Eopid())
+               {
+                       pexprIdent = pexprLeft;
+                       pexprConst = pexprRight;
+               }
+               else if (COperator::EopScalarIdent == 
pexprRight->Pop()->Eopid() &&
+                                COperator::EopScalarConst == 
pexprLeft->Pop()->Eopid())
+               {
+                       pexprIdent = pexprRight;
+                       pexprConst = pexprLeft;
+               }
+               else
+               {
+                       // ident = ident or const = const: not useful here
+                       continue;
+               }
+
+               // Skip a NULL constant: "col = NULL" never matches a row, so 
it must not
+               // drive pruning (substituting NULL into 
satisfies_hash_partition does not
+               // correspond to an equality match on a concrete key value).
+               if (colref == 
CScalarIdent::PopConvert(pexprIdent->Pop())->Pcr() &&
+                       
!CScalarConst::PopConvert(pexprConst->Pop())->GetDatum()->IsNull())
+               {
+                       return pexprConst;
+               }
+       }
+
+       return nullptr;
+}
+
+// Static pruning of a single HASH partition leaf, mirroring PostgreSQL hash
+// pruning by reusing satisfies_hash_partition(). The leaf's part constraint is
+// satisfies_hash_partition(parentoid, modulus, remainder, key...). We 
substitute
+// the equality constants for the partition-key columns (taken from the query
+// predicates) into that call and evaluate it with the constant-expression
+// evaluator. If the call evaluates to false, no row with those key values can
+// live in this leaf, so it is pruned. Any uncertainty (no equality on some 
key,
+// non-foldable constant, evaluator unavailable, non-bool result) keeps the 
leaf,
+// so pruning is always safe.
+BOOL
+CExpressionPreprocessor::FHashPartitionPruned(
+       CMemoryPool *mp, const IMDRelation *partrel, CColRefArray 
*pdrgpcrOutput,
+       ColRefToUlongMap *root_col_mapping, CExpressionArray 
*pdrgpexprConjuncts)
+{
+       IConstExprEvaluator *pceeval = COptCtxt::PoctxtFromTLS()->Pceeval();
+       if (nullptr == pceeval || !pceeval->FCanEvalExpressions())
+       {
+               return false;
+       }
+
+       CExpression *pexprCnstr = PexprPartConstraintFromChild(
+               mp, partrel, pdrgpcrOutput, root_col_mapping);
+       if (nullptr == pexprCnstr)
+       {
+               return false;
+       }
+
+       // The hash partition qual must be a satisfies_hash_partition() call 
with at
+       // least the (parentoid, modulus, remainder, key) arguments.
+       if (COperator::EopScalarFunc != pexprCnstr->Pop()->Eopid() ||
+               pexprCnstr->Arity() < 4)
+       {
+               pexprCnstr->Release();
+               return false;
+       }
+
+       // Rebuild the call, replacing each partition-key argument (a column 
ident)
+       // with the constant from a matching "key = const" predicate. If any key
+       // column lacks such an equality we cannot prune (PostgreSQL likewise 
needs
+       // equality on all hash key columns).
+       CExpressionArray *pdrgpexprArgs = GPOS_NEW(mp) CExpressionArray(mp);
+       BOOL fAllKeysBound = true;
+       const ULONG arity = pexprCnstr->Arity();
+       for (ULONG ul = 0; ul < arity; ++ul)
+       {
+               CExpression *pexprArg = (*pexprCnstr)[ul];
+               if (ul < 3)
+               {
+                       // parentoid, modulus, remainder: carry over unchanged
+                       pexprArg->AddRef();
+                       pdrgpexprArgs->Append(pexprArg);
+                       continue;
+               }
+
+               if (COperator::EopScalarIdent != pexprArg->Pop()->Eopid())
+               {
+                       fAllKeysBound = false;
+                       break;
+               }
+               const CColRef *colref = 
CScalarIdent::PopConvert(pexprArg->Pop())->Pcr();
+               CExpression *pexprConst =
+                       PexprColumnEqualityConst(pdrgpexprConjuncts, colref);
+               if (nullptr == pexprConst)
+               {
+                       fAllKeysBound = false;
+                       break;
+               }
+               pexprConst->AddRef();
+               pdrgpexprArgs->Append(pexprConst);
+       }
+
+       if (!fAllKeysBound)
+       {
+               pdrgpexprArgs->Release();
+               pexprCnstr->Release();
+               return false;
+       }
+
+       COperator *popFunc = pexprCnstr->Pop();
+       popFunc->AddRef();
+       CExpression *pexprTest = GPOS_NEW(mp) CExpression(mp, popFunc, 
pdrgpexprArgs);
+       pexprCnstr->Release();
+
+       // satisfies_hash_partition() over constants should always fold to a 
boolean,
+       // but if evaluation fails for any reason we must not abort 
optimization for
+       // the whole query: reset the error and conservatively keep the 
partition
+       // (pexprResult stays null, so we fall through to return false).
+       CExpression *pexprResult = nullptr;
+       GPOS_TRY
+       {
+               pexprResult = pceeval->PexprEval(pexprTest);
+       }
+       GPOS_CATCH_EX(ex)
+       {
+               GPOS_RESET_EX;
+       }
+       GPOS_CATCH_END;
+       pexprTest->Release();
+
+       BOOL fPruned = false;
+       if (nullptr != pexprResult &&
+               COperator::EopScalarConst == pexprResult->Pop()->Eopid())
+       {
+               IDatum *datum = 
CScalarConst::PopConvert(pexprResult->Pop())->GetDatum();
+               IDatumBool *pdatumBool = dynamic_cast<IDatumBool *>(datum);
+               if (nullptr != pdatumBool && !pdatumBool->IsNull() &&
+                       !pdatumBool->GetValue())
+               {
+                       // satisfies_hash_partition() is false => this leaf 
cannot hold any
+                       // row with the given key values => prune it.
+                       fPruned = true;
+               }
+       }
+       CRefCount::SafeRelease(pexprResult);
+
+       return fPruned;
+}
+
 // Transpose a select over a project
 //
 // This preprocessing step enables additional opportunities for predicate push
diff --git a/src/backend/gporca/libnaucrates/include/naucrates/md/IMDRelation.h 
b/src/backend/gporca/libnaucrates/include/naucrates/md/IMDRelation.h
index b5e7a8f22ee..f0b5c74f4c0 100644
--- a/src/backend/gporca/libnaucrates/include/naucrates/md/IMDRelation.h
+++ b/src/backend/gporca/libnaucrates/include/naucrates/md/IMDRelation.h
@@ -75,7 +75,8 @@ public:
        enum Erelpartitiontype
        {
                ErelpartitionRange = 'r',
-               ErelpartitionList = 'l'
+               ErelpartitionList = 'l',
+               ErelpartitionHash = 'h'
        };
 
 protected:
diff --git a/src/test/regress/expected/direct_dispatch_optimizer.out 
b/src/test/regress/expected/direct_dispatch_optimizer.out
index 2da45e7955a..57e499bb091 100644
--- a/src/test/regress/expected/direct_dispatch_optimizer.out
+++ b/src/test/regress/expected/direct_dispatch_optimizer.out
@@ -1675,15 +1675,11 @@ explain select * from t_14887 where a = 'a'::text;
                                     QUERY PLAN                                 
   
 
----------------------------------------------------------------------------------
  Gather Motion 3:1  (slice1; segments: 3)  (cost=0.00..764.38 rows=158 
width=32)
-   ->  Append  (cost=0.00..762.26 rows=53 width=32)
-         ->  Seq Scan on t0_14887 t_14887_1  (cost=0.00..254.00 rows=18 
width=32)
-               Filter: ((a)::text = 'a'::text)
-         ->  Seq Scan on t1_14887 t_14887_2  (cost=0.00..254.00 rows=18 
width=32)
-               Filter: ((a)::text = 'a'::text)
-         ->  Seq Scan on t2_14887 t_14887_3  (cost=0.00..254.00 rows=18 
width=32)
-               Filter: ((a)::text = 'a'::text)
- Optimizer: Postgres query optimizer
-(9 rows)
+   ->  Dynamic Seq Scan on t_14887
+         Number of partitions to scan: 3 (out of 3)
+         Filter: ((a)::text = 'a'::text)
+ Optimizer: Pivotal Optimizer (GPORCA)
+(4 rows)
 
 begin;
 drop table if exists direct_test;
diff --git a/src/test/regress/expected/orca_static_pruning_optimizer.out 
b/src/test/regress/expected/orca_static_pruning_optimizer.out
index 522f0dc17f7..9640d004ef0 100644
--- a/src/test/regress/expected/orca_static_pruning_optimizer.out
+++ b/src/test/regress/expected/orca_static_pruning_optimizer.out
@@ -160,21 +160,18 @@ WHERE b = 42
 $query$ AS qry \gset
 EXPLAIN (COSTS OFF, VERBOSE)
 :qry ;
-INFO:  GPORCA failed to produce a plan, falling back to Postgres-based planner
-DETAIL:  Falling back to Postgres-based planner because GPORCA does not 
support the following feature: hash partitioning
                   QUERY PLAN                  
 ----------------------------------------------
  Gather Motion 3:1  (slice1; segments: 3)
-   Output: hp.a, hp.b
-   ->  Seq Scan on orca_static_pruning.hp0 hp
-         Output: hp.a, hp.b
+   Output: a, b
+   ->  Dynamic Seq Scan on orca_static_pruning.hp
+         Output: a, b
+         Number of partitions to scan: 1 (out of 2)
          Filter: (hp.b = 42)
- Optimizer: Postgres query optimizer
-(6 rows)
+ Optimizer: Pivotal Optimizer (GPORCA)
+(8 rows)
 
 :qry ;
-INFO:  GPORCA failed to produce a plan, falling back to Postgres-based planner
-DETAIL:  Falling back to Postgres-based planner because GPORCA does not 
support the following feature: hash partitioning
  a | b  
 ---+----
  0 | 42
diff --git a/src/test/regress/expected/partition_prune_append.out 
b/src/test/regress/expected/partition_prune_append.out
index 20730d90d0f..75ae62577d8 100644
--- a/src/test/regress/expected/partition_prune_append.out
+++ b/src/test/regress/expected/partition_prune_append.out
@@ -3502,30 +3502,32 @@ explain (costs off) select * from pph_arrpart where a = 
'{1}';
                  QUERY PLAN                 
 --------------------------------------------
  Gather Motion 1:1  (slice1; segments: 1)
-   ->  Seq Scan on pph_arrpart2 pph_arrpart
-         Filter: (a = '{1}'::integer[])
- Optimizer: Postgres query optimizer
-(4 rows)
+   ->  Append
+         ->  Seq Scan on pph_arrpart2
+               Filter: (a = '{1}'::integer[])
+ Optimizer: Pivotal Optimizer (GPORCA)
+(5 rows)
 
 explain (costs off) select * from pph_arrpart where a = '{1, 2}';
                  QUERY PLAN                 
 --------------------------------------------
  Gather Motion 1:1  (slice1; segments: 1)
-   ->  Seq Scan on pph_arrpart1 pph_arrpart
-         Filter: (a = '{1,2}'::integer[])
- Optimizer: Postgres query optimizer
-(4 rows)
+   ->  Append
+         ->  Seq Scan on pph_arrpart1
+               Filter: (a = '{1,2}'::integer[])
+ Optimizer: Pivotal Optimizer (GPORCA)
+(5 rows)
 
 explain (costs off) select * from pph_arrpart where a in ('{4, 5}', '{1}');
                                  QUERY PLAN                                 
 ----------------------------------------------------------------------------
- Gather Motion 2:1  (slice1; segments: 2)
+ Gather Motion 3:1  (slice1; segments: 3)
    ->  Append
-         ->  Seq Scan on pph_arrpart1 pph_arrpart_1
+         ->  Seq Scan on pph_arrpart1
                Filter: ((a = '{4,5}'::integer[]) OR (a = '{1}'::integer[]))
-         ->  Seq Scan on pph_arrpart2 pph_arrpart_2
+         ->  Seq Scan on pph_arrpart2
                Filter: ((a = '{4,5}'::integer[]) OR (a = '{1}'::integer[]))
- Optimizer: Postgres query optimizer
+ Optimizer: Pivotal Optimizer (GPORCA)
 (7 rows)
 
 drop table pph_arrpart;
diff --git a/src/test/regress/expected/partition_prune_optimizer.out 
b/src/test/regress/expected/partition_prune_optimizer.out
index 72caa769ea3..d3fdaa3493f 100644
--- a/src/test/regress/expected/partition_prune_optimizer.out
+++ b/src/test/regress/expected/partition_prune_optimizer.out
@@ -3949,31 +3949,31 @@ explain (costs off) select * from pph_arrpart where a = 
'{1}';
                  QUERY PLAN                 
 --------------------------------------------
  Gather Motion 1:1  (slice1; segments: 1)
-   ->  Seq Scan on pph_arrpart2 pph_arrpart
+   ->  Dynamic Seq Scan on pph_arrpart
+         Number of partitions to scan: 1 (out of 2)
          Filter: (a = '{1}'::integer[])
- Optimizer: Postgres query optimizer
-(4 rows)
+ Optimizer: Pivotal Optimizer (GPORCA)
+(5 rows)
 
 explain (costs off) select * from pph_arrpart where a = '{1, 2}';
                  QUERY PLAN                 
 --------------------------------------------
  Gather Motion 1:1  (slice1; segments: 1)
-   ->  Seq Scan on pph_arrpart1 pph_arrpart
+   ->  Dynamic Seq Scan on pph_arrpart
+         Number of partitions to scan: 1 (out of 2)
          Filter: (a = '{1,2}'::integer[])
- Optimizer: Postgres query optimizer
-(4 rows)
+ Optimizer: Pivotal Optimizer (GPORCA)
+(5 rows)
 
 explain (costs off) select * from pph_arrpart where a in ('{4, 5}', '{1}');
                                  QUERY PLAN                                 
 ----------------------------------------------------------------------------
- Gather Motion 2:1  (slice1; segments: 2)
-   ->  Append
-         ->  Seq Scan on pph_arrpart1 pph_arrpart_1
-               Filter: ((a = '{4,5}'::integer[]) OR (a = '{1}'::integer[]))
-         ->  Seq Scan on pph_arrpart2 pph_arrpart_2
-               Filter: ((a = '{4,5}'::integer[]) OR (a = '{1}'::integer[]))
- Optimizer: Postgres query optimizer
-(7 rows)
+ Gather Motion 3:1  (slice1; segments: 3)
+   ->  Dynamic Seq Scan on pph_arrpart
+         Number of partitions to scan: 2 (out of 2)
+         Filter: ((a = '{4,5}'::integer[]) OR (a = '{1}'::integer[]))
+ Optimizer: Pivotal Optimizer (GPORCA)
+(5 rows)
 
 drop table pph_arrpart;
 -- enum type list partition key
diff --git a/src/test/regress/expected/qp_indexscan_optimizer.out 
b/src/test/regress/expected/qp_indexscan_optimizer.out
index 8d187de1d73..ae5405183e9 100644
--- a/src/test/regress/expected/qp_indexscan_optimizer.out
+++ b/src/test/regress/expected/qp_indexscan_optimizer.out
@@ -2403,14 +2403,13 @@ explain(costs off) select * from tbl_hash order by col1 
limit 3;
 -----------------------------------------------------------------------
  Limit
    ->  Gather Motion 3:1  (slice1; segments: 3)
-         Merge Key: tbl_hash.col1
-         ->  Limit
-               ->  Merge Append
-                     Sort Key: tbl_hash.col1
-                     ->  Index Scan using p1_col1_idx on p1 tbl_hash_1
-                     ->  Index Scan using p2_col1_idx on p2 tbl_hash_2
- Optimizer: Postgres query optimizer
-(9 rows)
+         Merge Key: col1
+         ->  Sort
+               Sort Key: col1
+               ->  Dynamic Seq Scan on tbl_hash
+                     Number of partitions to scan: 2 (out of 2)
+ Optimizer: Pivotal Optimizer (GPORCA)
+(8 rows)
 
 -- Clean Up
 DROP TABLE tbl_hash;


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to