Hi Matthias!
Thanks a lot for your review.
Please add commit descriptions; the current patches don't have much
other than their subjects which don't describe any rationale for the
changes applied.
Done.
0001: LGTM
0002:
1. radix_sort_trigrams_signed has an `int count` argument, but the
caller trigram_qsort uses size_t.
Changed count argument to size_t.
2. This increases the memory requirements of trigram_qsort by a huge margin.
Could you change the radixsort to operate in-place, so that the new
buffer is not needed?
Trigram radix sort is only called from generate_trgm() and
generate_wildcard_trgm() which are both used to extract the unique
trigrams from a _single_ string value.
We don't radix sort trigrams of multiple string values. Hence, the
maximum increase in memory, while building the GIN index, is in the size
of the longest string encountered, not in the number of rows.
Given that in-place radix sort is a lot more complex and slower, I think
the current tradeoff is fine.
Beyond that, other GIN index functions also allocate extra memory on a
per-value basis which shows that this should not be a problem in
practice (e.g. ginarrayextract(), gin_extract_value_trgm(), ...).
3. The implementation for trigram_qsort_unsigned has not been adjusted
nor replaced, and so keeps the same old performance that
trigram_qsort_signed had.
Please make sure to also adjust that implementation.
Done.
I laid out the code such that the compiler has the possibility to fully
inline both variants to get rid of the extra code in radix_key() for
flipping the sign bit in the unsigned case. But even if it doesn't,
radix_key() can be branchless and the performance is anyway dominated by
memory traffic.
0003:
1. ginInsertBAEntry leaks the copied key Datum if the type is by-ref
and already present in the accumulator.
Good catch!
They weren't leaked for good but would have gotten cleaned up with the
next batch. But they would have unnecessarily increased the memory
footprint. Especially, as the number of unique trigrams is typically low
and hence we often hit the "already present in the accumulator" path.
I think specifying HASH_KEYCOPY to only copying the Datum key if an
entry is already present is more appropriate.
I cannot use HASH_KEYCOPY because I'm using simplehash.h. I changed it
to use the original key for doing the lookup and only create the copy in
the !found path. This is also how the old code did it. Changing the key
of an already inserted hashmap entry is save because we overwrite it
with a copy. So the hash function will keep returning the same position
for it.
2. The comment for ginInsertBAEntries was completely removed, rather
than updated to the new workings.
Yes, because the comment no longer applies. All of it was specific to
using a red-black tree. ginInsertBAEntries() is now merely syntactic
sugar. I thought about removing the function completely but it's used in
a couple of places and without it, the call sites would get slightly
more ugly.
Would you like to see anything specific in the comment?
Attached is the updated patch, rebased on latest master. With 0003,
rbtree.h/.c are completely unused and could be removed as well,
including the test code.
--
David Geier
From cf9dd48306cf0399ea45d540b55aacf744210cef Mon Sep 17 00:00:00 2001
From: David Geier <[email protected]>
Date: Mon, 10 Nov 2025 15:40:11 +0100
Subject: [PATCH v10 1/3] Use branchless comparisons in btint4cmp and btint8cmp
Use the common pg_cmp_s32() and pg_cmp_s64() helpers to implement the
built-in B-tree comparison functions for int4 and int8.
The previous implementations used conditional branches to distinguish
less-than, equal, and greater-than values. The common comparison helpers
perform the same three-way comparison without data-dependent branches,
which can improve performance for workloads involving frequent integer
comparisons while preserving the required comparator result semantics.
btint4cmp() and btint8cmp() are PostgreSQL-callable functions invoked
through the function manager. They are not inlined at their call sites,
so replacing the original conditional implementation does not prevent a
compiler from optimizing an inline comparison in contexts where it can
see and better optimize the surrounding code. In other words, this
change affects the function-manager call path without imposing a
performance regression on callers for which the comparison could
otherwise have been inlined.
The comparison helpers also provide the appropriate handling for the
full ranges of int32 and int64 values without relying on subtraction,
which could overflow for values near the type limits.
---
src/backend/access/nbtree/nbtcompare.c | 15 +++------------
1 file changed, 3 insertions(+), 12 deletions(-)
diff --git a/src/backend/access/nbtree/nbtcompare.c
b/src/backend/access/nbtree/nbtcompare.c
index 4e3a3a0f7ce..80dec200a3d 100644
--- a/src/backend/access/nbtree/nbtcompare.c
+++ b/src/backend/access/nbtree/nbtcompare.c
@@ -61,6 +61,7 @@
#include "utils/fmgrprotos.h"
#include "utils/skipsupport.h"
#include "utils/sortsupport.h"
+#include "common/int.h"
#ifdef STRESS_SORT_INT_MIN
#define A_LESS_THAN_B INT_MIN
@@ -194,12 +195,7 @@ btint4cmp(PG_FUNCTION_ARGS)
int32 a = PG_GETARG_INT32(0);
int32 b = PG_GETARG_INT32(1);
- if (a > b)
- PG_RETURN_INT32(A_GREATER_THAN_B);
- else if (a == b)
- PG_RETURN_INT32(0);
- else
- PG_RETURN_INT32(A_LESS_THAN_B);
+ PG_RETURN_INT32(pg_cmp_s32(a, b));
}
Datum
@@ -262,12 +258,7 @@ btint8cmp(PG_FUNCTION_ARGS)
int64 a = PG_GETARG_INT64(0);
int64 b = PG_GETARG_INT64(1);
- if (a > b)
- PG_RETURN_INT32(A_GREATER_THAN_B);
- else if (a == b)
- PG_RETURN_INT32(0);
- else
- PG_RETURN_INT32(A_LESS_THAN_B);
+ PG_RETURN_INT32(pg_cmp_s64(a, b));
}
Datum
--
2.50.1 (Apple Git-155)
From 24f53244f6e8f44957e7cb1ff0e6cdd920aaab19 Mon Sep 17 00:00:00 2001
From: David Geier <[email protected]>
Date: Tue, 11 Nov 2025 13:18:59 +0100
Subject: [PATCH v10 2/3] Use radix sort to extract trigrams
Replace the comparison-based sort used by generate_trgm() and
generate_wildcard_trgm() with a three-pass radix sort.
Trigrams consist of three bytes, so their keys have a fixed and very
small width. A radix sort can therefore order them in linear time with
respect to the number of trigrams, avoiding the repeated comparator calls
and recursive partitioning performed by qsort.
The implementation preserves the existing behavior on platforms where
char is signed by flipping the most significant bit before sorting.
The radix sort requires a temporary buffer, increasing the memory
footprint while trigrams are being extracted from an input string.
However, the sort is performed separately for each string and is never
applied across multiple strings at once. Consequently, the additional
memory is limited to the processing of the current string and should
have a negligible effect on the overall memory footprint of a GIN index
build.
The resulting order remains compatible with trigram deduplication and
does not change the generated trigram sets.
---
contrib/pg_trgm/trgm_op.c | 78 ++++++++++++++++++++++++++++-----------
1 file changed, 57 insertions(+), 21 deletions(-)
diff --git a/contrib/pg_trgm/trgm_op.c b/contrib/pg_trgm/trgm_op.c
index 22bcc3c3361..bfda0df9fd4 100644
--- a/contrib/pg_trgm/trgm_op.c
+++ b/contrib/pg_trgm/trgm_op.c
@@ -226,33 +226,69 @@ CMPTRGM_CHOOSE(const void *a, const void *b)
return CMPTRGM(a, b);
}
-#define ST_SORT trigram_qsort_signed
-#define ST_ELEMENT_TYPE_VOID
-#define ST_COMPARE(a, b) CMPTRGM_SIGNED(a, b)
-#define ST_SCOPE static
-#define ST_DEFINE
-#define ST_DECLARE
-#include "lib/sort_template.h"
-
-#define ST_SORT trigram_qsort_unsigned
-#define ST_ELEMENT_TYPE_VOID
-#define ST_COMPARE(a, b) CMPTRGM_UNSIGNED(a, b)
-#define ST_SCOPE static
-#define ST_DEFINE
-#define ST_DECLARE
-#include "lib/sort_template.h"
+/*
+ * Needed to properly handle negative numbers in case char is signed.
+ */
+static inline unsigned char
+radix_key(char x, bool char_is_signed)
+{
+ return char_is_signed ? x ^ 0x80 : x;
+}
+
+static inline void
+trigram_radix_sort_with_signedness(trgm *trg, size_t count, bool
char_is_signed)
+{
+ trgm *buffer = palloc_array(trgm, count);
+ trgm *starts[256];
+ trgm *from = trg;
+ trgm *to = buffer;
+ size_t freqs[3][256];
+
+ /*
+ * Compute frequencies to partition the buffer.
+ */
+ memset(freqs, 0, sizeof(freqs));
+
+ for (size_t i = 0; i < count; i++)
+ for (size_t j = 0; j < 3; j++)
+ freqs[j][radix_key(trg[i][j], char_is_signed)]++;
+
+ /*
+ * Do the sorting. Start with last character because that's the "LSB"
+ * in a trigram. Avoid unnecessary copies by ping-ponging between the
buffers.
+ */
+ for (int i = 2; i >= 0; i--)
+ {
+ trgm *old_from = from;
+ trgm *next = to;
+
+ for (size_t j = 0; j < 256; j++)
+ {
+ starts[j] = next;
+ next += freqs[i][j];
+ }
+
+ for (size_t j = 0; j < count; j++)
+ memcpy(starts[radix_key(from[j][i], char_is_signed)]++,
from[j], sizeof(trgm));
+
+ from = to;
+ to = old_from;
+ }
+
+ memcpy(trg, buffer, sizeof(trgm) * count);
+ pfree(buffer);
+}
/* Sort an array of trigrams, handling signedness correctly */
static void
-trigram_qsort(trgm *array, size_t n)
+trigram_radix_sort(trgm *array, size_t n)
{
if (GetDefaultCharSignedness())
- trigram_qsort_signed(array, n, sizeof(trgm));
+ trigram_radix_sort_with_signedness(array, n, true);
else
- trigram_qsort_unsigned(array, n, sizeof(trgm));
+ trigram_radix_sort_with_signedness(array, n, false);
}
-
/*
* Compare two trigrams for equality. This has the same signature as
* comparison functions used for sorting, so that this can be used with
@@ -612,7 +648,7 @@ generate_trgm(char *str, int slen)
*/
if (len > 1)
{
- trigram_qsort(GETARR(trg), len);
+ trigram_radix_sort(GETARR(trg), len);
len = trigram_qunique(GETARR(trg), len);
}
@@ -1143,7 +1179,7 @@ generate_wildcard_trgm(const char *str, int slen)
len = arr.length;
if (len > 1)
{
- trigram_qsort(GETARR(trg), len);
+ trigram_radix_sort(GETARR(trg), len);
len = trigram_qunique(GETARR(trg), len);
}
--
2.50.1 (Apple Git-155)
From d24a1ff663fcaf0d8adf3661dbc65efa3dddf403 Mon Sep 17 00:00:00 2001
From: David Geier <[email protected]>
Date: Wed, 22 Apr 2026 14:00:40 +0200
Subject: [PATCH v10 3/3] Replace GIN build accumulator RB-tree with a hashmap
Replace the red-black tree used by ginInsertBAEntries() with a hash table
based on simplehash.h. This keeps the same basic accumulation strategy
as the previous implementation while changing key lookup and deduplication
from O(log(num_unique_keys)) tree operations to expected O(1) hash-table
operations. As a result, the overall complexity changes from
O(num_total_keys * log(num_unique_keys))
to
O(num_total_keys + num_unique_keys * log(num_unique_keys))
The latter is preferable for the usual case, where the number of unique
keys is much smaller than the number of rows. Even when most or all keys
are unique, the RB-tree rebalancing operations are sufficiently
expensive that the theoretical worst-case advantage of the tree does not
necessarily translate into better runtime.
The item pointer lists associated with each key continue to be sorted
before the accumulated entries are emitted. Use sort_template.h for this
sorting instead of qsort(). The distinct hash entries are copied into an
array and sorted using the existing GIN key comparison function so that
the output order remains unchanged.
ginInsertBAEntries() got much simpler as well. The previous implementation
inserted entries in an order intended to minimize red-black tree rebalancing
for sorted inputs. With a hash table, inserting the entries in their original
order is preferable because repeated keys are more likely to be found in the
same hash-table entry.
Use datumIsEqual() and datum_image_hash() for normal GIN keys. This is
consistent with the current GIN implementation, which does not correctly
support non-deterministic collations or types for which logically equal
values are not image-equal.
Because parallel GIN index builds also use ginInsertBAEntries(), this
change improves the accumulation phase of parallel index builds as well.
---
src/backend/access/gin/ginbulk.c | 339 +++++++++++++++----------------
src/include/access/gin_private.h | 24 +--
src/tools/pgindent/typedefs.list | 4 +-
3 files changed, 175 insertions(+), 192 deletions(-)
diff --git a/src/backend/access/gin/ginbulk.c b/src/backend/access/gin/ginbulk.c
index 85865b39105..77c12ba932b 100644
--- a/src/backend/access/gin/ginbulk.c
+++ b/src/backend/access/gin/ginbulk.c
@@ -17,107 +17,118 @@
#include <limits.h>
#include "access/gin_private.h"
+#include "common/hashfn.h"
#include "utils/datum.h"
#include "utils/memutils.h"
+#define DEF_NENTRY 2048 /* Initial hash table size */
+#define DEF_ITEMS_PER_KEY 8 /* Initial ItemPointer array
size per key */
-#define DEF_NENTRY 2048 /* GinEntryAccumulator allocation
quantum */
-#define DEF_NPTR 5 /* ItemPointer initial
allocation quantum */
-
-
-/* Combiner function for rbtree.c */
-static void
-ginCombineData(RBTNode *existing, const RBTNode *newdata, void *arg)
+typedef struct GinHashKey
{
- GinEntryAccumulator *eo = (GinEntryAccumulator *) existing;
- const GinEntryAccumulator *en = (const GinEntryAccumulator *) newdata;
- BuildAccumulator *accum = (BuildAccumulator *) arg;
+ OffsetNumber attnum;
+ GinNullCategory category;
+ Datum key;
+} GinHashKey;
- /*
- * Note this code assumes that newdata contains only one itempointer.
- */
- if (eo->count >= eo->maxcount)
- {
- if (eo->maxcount > INT_MAX)
- ereport(ERROR,
-
(errcode(ERRCODE_PROGRAM_LIMIT_EXCEEDED),
- errmsg("posting list is too long"),
- errhint("Reduce
\"maintenance_work_mem\".")));
+typedef struct GinHashEntry
+{
+ GinHashKey hashkey;
+ uint32 hash;
+ char status;
+ ItemPointerData *items;
+ uint32 numItems;
+ uint32 allocatedItems;
+} GinHashEntry;
+
+typedef struct GinSortEntry
+{
+ GinHashKey hashkey;
+ ItemPointerData *items;
+ uint32 numItems;
+} GinSortEntry;
+
+static uint32 gin_hash_key(struct ginbuild_hash *tb, GinHashKey *key);
+static bool gin_equal_key(struct ginbuild_hash *tb, GinHashKey *a, GinHashKey
*b);
+
+#define SH_PREFIX ginbuild
+#define SH_ELEMENT_TYPE GinHashEntry
+#define SH_KEY_TYPE GinHashKey
+#define SH_KEY hashkey
+#define SH_HASH_KEY(tb, key) gin_hash_key(tb, &key)
+#define SH_EQUAL(tb, a, b) gin_equal_key(tb, &a, &b)
+#define SH_SCOPE static inline
+#define SH_STORE_HASH
+#define SH_GET_HASH(tb, a) (a)->hash
+#define SH_DEFINE
+#define SH_DECLARE
+#include "lib/simplehash.h"
+
+static uint32
+gin_hash_key(struct ginbuild_hash *tb, GinHashKey *key)
+{
+ BuildAccumulator *accum = (BuildAccumulator *) tb->private_data;
+ uint32 hash;
- accum->allocatedMemory -= GetMemoryChunkSpace(eo->list);
- eo->maxcount *= 2;
- eo->list = (ItemPointerData *)
- repalloc_huge(eo->list, sizeof(ItemPointerData) *
eo->maxcount);
- accum->allocatedMemory += GetMemoryChunkSpace(eo->list);
- }
+ hash = hash_combine(0, murmurhash32((uint32) key->attnum));
+ hash = hash_combine(hash, murmurhash32((uint32) key->category));
- /* If item pointers are not ordered, they will need to be sorted later
*/
- if (eo->shouldSort == false)
+ if (key->category == GIN_CAT_NORM_KEY)
{
- int res;
-
- res = ginCompareItemPointers(eo->list + eo->count - 1,
en->list);
- Assert(res != 0);
+ CompactAttribute *att;
- if (res > 0)
- eo->shouldSort = true;
+ att = TupleDescCompactAttr(accum->ginstate->origTupdesc,
key->attnum - 1);
+ hash = hash_combine(hash, datum_image_hash(key->key,
att->attbyval, att->attlen));
}
- eo->list[eo->count] = en->list[0];
- eo->count++;
+ return hash;
}
-/* Comparator function for rbtree.c */
-static int
-cmpEntryAccumulator(const RBTNode *a, const RBTNode *b, void *arg)
+static bool
+gin_equal_key(struct ginbuild_hash *tb, GinHashKey *a, GinHashKey *b)
{
- const GinEntryAccumulator *ea = (const GinEntryAccumulator *) a;
- const GinEntryAccumulator *eb = (const GinEntryAccumulator *) b;
- BuildAccumulator *accum = (BuildAccumulator *) arg;
-
- return ginCompareAttEntries(accum->ginstate,
- ea->attnum,
ea->key, ea->category,
- eb->attnum,
eb->key, eb->category);
-}
+ BuildAccumulator *accum = (BuildAccumulator *) tb->private_data;
+ CompactAttribute *att;
-/* Allocator function for rbtree.c */
-static RBTNode *
-ginAllocEntryAccumulator(void *arg)
-{
- BuildAccumulator *accum = (BuildAccumulator *) arg;
- GinEntryAccumulator *ea;
+ if (a->attnum != b->attnum)
+ return false;
+ if (a->category != b->category)
+ return false;
+ if (a->category != GIN_CAT_NORM_KEY)
+ return true;
/*
- * Allocate memory by rather big chunks to decrease overhead. We have
no
- * need to reclaim RBTNodes individually, so this costs nothing.
+ * Compare the actual key values using image equality.
+ * This is correct because we don't want to deduplicate at this point.
*/
- if (accum->entryallocator == NULL || accum->eas_used >= DEF_NENTRY)
- {
- accum->entryallocator = palloc_array(GinEntryAccumulator,
DEF_NENTRY);
- accum->allocatedMemory +=
GetMemoryChunkSpace(accum->entryallocator);
- accum->eas_used = 0;
- }
-
- /* Allocate new RBTNode from current chunk */
- ea = accum->entryallocator + accum->eas_used;
- accum->eas_used++;
-
- return (RBTNode *) ea;
+ att = TupleDescCompactAttr(accum->ginstate->origTupdesc, a->attnum - 1);
+ return datumIsEqual(a->key, b->key, att->attbyval, att->attlen);
}
+#define ST_SORT sort_itempointers
+#define ST_ELEMENT_TYPE ItemPointerData
+#define ST_COMPARE(a, b) ginCompareItemPointers(a, b)
+#define ST_SCOPE static
+#define ST_DEFINE
+#include "lib/sort_template.h"
+
+#define ST_SORT sort_keys
+#define ST_ELEMENT_TYPE GinSortEntry
+#define ST_COMPARE_ARG_TYPE GinState
+#define ST_COMPARE(a, b, state) ginCompareAttEntries(state, a->hashkey.attnum,
a->hashkey.key, a->hashkey.category, b->hashkey.attnum, b->hashkey.key,
b->hashkey.category)
+#define ST_SCOPE static
+#define ST_DEFINE
+#include "lib/sort_template.h"
+
void
ginInitBA(BuildAccumulator *accum)
{
/* accum->ginstate is intentionally not set here */
- accum->allocatedMemory = 0;
- accum->entryallocator = NULL;
- accum->eas_used = 0;
- accum->tree = rbt_create(sizeof(GinEntryAccumulator),
- cmpEntryAccumulator,
- ginCombineData,
-
ginAllocEntryAccumulator,
- NULL, /* no freefunc
needed */
- accum);
+ accum->hash = ginbuild_create(CurrentMemoryContext, DEF_NENTRY, accum);
+ accum->allocatedMemory = accum->hash->size * sizeof(GinHashEntry);
+ accum->sorted_entries = NULL;
+ accum->num_entries = 0;
+ accum->current_pos = 0;
}
/*
@@ -142,124 +153,113 @@ getDatumCopy(BuildAccumulator *accum, OffsetNumber
attnum, Datum value)
}
/*
- * Find/store one entry from indexed value.
+ * Insert one entry into the hash map.
+ * If the key already exists, append to its ItemPointer array.
+ * Otherwise, create a new hash entry with a new ItemPointer array.
*/
static void
ginInsertBAEntry(BuildAccumulator *accum,
ItemPointer heapptr, OffsetNumber attnum,
Datum key, GinNullCategory category)
{
- GinEntryAccumulator eatmp;
- GinEntryAccumulator *ea;
- bool isNew;
+ GinHashKey hashkey;
+ GinHashEntry *entry;
+ bool found;
+ uint64 oldsize;
- /*
- * For the moment, fill only the fields of eatmp that will be looked at
by
- * cmpEntryAccumulator or ginCombineData.
- */
- eatmp.attnum = attnum;
- eatmp.key = key;
- eatmp.category = category;
- /* temporarily set up single-entry itempointer list */
- eatmp.list = heapptr;
+ hashkey.attnum = attnum;
+ hashkey.category = category;
+ hashkey.key = key;
- ea = (GinEntryAccumulator *) rbt_insert(accum->tree, (RBTNode *) &eatmp,
-
&isNew);
+ oldsize = accum->hash->size;
+ entry = ginbuild_insert(accum->hash, hashkey, &found);
- if (isNew)
+ if (!found)
{
/*
- * Finish initializing new tree entry, including making
permanent
- * copies of the datum (if it's not null) and itempointer.
+ * Finish initializing new hashmap entry including making a
permanent
+ * copy of the key.
*/
if (category == GIN_CAT_NORM_KEY)
- ea->key = getDatumCopy(accum, attnum, key);
- ea->maxcount = DEF_NPTR;
- ea->count = 1;
- ea->shouldSort = false;
- ea->list = palloc_array(ItemPointerData, DEF_NPTR);
- ea->list[0] = *heapptr;
- accum->allocatedMemory += GetMemoryChunkSpace(ea->list);
+ entry->hashkey.key = getDatumCopy(accum, attnum, key);
+
+ entry->items = palloc_array(ItemPointerData, DEF_ITEMS_PER_KEY);
+ entry->numItems = 0;
+ entry->allocatedItems = DEF_ITEMS_PER_KEY;
+ accum->allocatedMemory += (accum->hash->size - oldsize) *
sizeof(GinHashEntry);
+ accum->allocatedMemory += GetMemoryChunkSpace(entry->items);
}
- else
+
+ if (entry->numItems >= entry->allocatedItems)
{
- /*
- * ginCombineData did everything needed.
- */
+ uint32 new_allocated;
+
+ if (entry->allocatedItems > UINT32_MAX / 2)
+ ereport(ERROR,
+
(errcode(ERRCODE_PROGRAM_LIMIT_EXCEEDED),
+ errmsg("too many GIN item pointers for
a single key"),
+ errhint("Reduce
\"maintenance_work_mem\".")));
+
+ accum->allocatedMemory -= GetMemoryChunkSpace(entry->items);
+ new_allocated = entry->allocatedItems * 2;
+ entry->items = repalloc_huge(entry->items,
mul_size(sizeof(ItemPointerData), new_allocated));
+ entry->allocatedItems = new_allocated;
+ accum->allocatedMemory += GetMemoryChunkSpace(entry->items);
}
+
+ entry->items[entry->numItems++] = *heapptr;
}
-/*
- * Insert the entries for one heap pointer.
- *
- * Since the entries are being inserted into a balanced binary tree, you
- * might think that the order of insertion wouldn't be critical, but it turns
- * out that inserting the entries in sorted order results in a lot of
- * rebalancing operations and is slow. To prevent this, we attempt to insert
- * the nodes in an order that will produce a nearly-balanced tree if the input
- * is in fact sorted.
- *
- * We do this as follows. First, we imagine that we have an array whose size
- * is the smallest power of two greater than or equal to the actual array
- * size. Second, we insert the middle entry of our virtual array into the
- * tree; then, we insert the middles of each half of our virtual array, then
- * middles of quarters, etc.
- */
void
ginInsertBAEntries(BuildAccumulator *accum,
ItemPointer heapptr, OffsetNumber attnum,
Datum *entries, GinNullCategory *categories,
int32 nentries)
{
- uint32 step = nentries;
-
if (nentries <= 0)
return;
Assert(ItemPointerIsValid(heapptr) && attnum >= FirstOffsetNumber);
- /*
- * step will contain largest power of 2 and <= nentries
- */
- step |= (step >> 1);
- step |= (step >> 2);
- step |= (step >> 4);
- step |= (step >> 8);
- step |= (step >> 16);
- step >>= 1;
- step++;
-
- while (step > 0)
- {
- int i;
-
- for (i = step - 1; i < nentries && i >= 0; i += step << 1 /* *2
*/ )
- ginInsertBAEntry(accum, heapptr, attnum,
- entries[i],
categories[i]);
-
- step >>= 1; /* /2 */
- }
+ for (int i = 0; i < nentries; i++)
+ ginInsertBAEntry(accum, heapptr, attnum, entries[i],
categories[i]);
}
-static int
-qsortCompareItemPointers(const void *a, const void *b)
-{
- int res = ginCompareItemPointers((const
ItemPointerData *) a, (const ItemPointerData *) b);
-
- /* Assert that there are no equal item pointers being sorted */
- Assert(res != 0);
- return res;
-}
-
-/* Prepare to read out the rbtree contents using ginGetBAEntry */
+/* Prepare to read out the hash table contents using ginGetBAEntry */
void
ginBeginBAScan(BuildAccumulator *accum)
{
- rbt_begin_iterate(accum->tree, LeftRightWalk, &accum->tree_walk);
+ ginbuild_iterator iter;
+ GinHashEntry *entry;
+ uint32 i = 0;
+
+ accum->num_entries = accum->hash->members;
+ accum->current_pos = 0;
+
+ if (accum->num_entries == 0)
+ return;
+
+ accum->sorted_entries = palloc_array(GinSortEntry, accum->num_entries);
+ ginbuild_start_iterate(accum->hash, &iter);
+
+ while ((entry = ginbuild_iterate(accum->hash, &iter)) != NULL)
+ {
+ GinSortEntry *se = &accum->sorted_entries[i];
+ sort_itempointers(entry->items, entry->numItems);
+
+ se->hashkey = entry->hashkey;
+ se->items = entry->items;
+ se->numItems = entry->numItems;
+ i++;
+ }
+
+ Assert(i == accum->num_entries);
+ sort_keys(accum->sorted_entries, accum->num_entries, accum->ginstate);
+ accum->current_pos = 0;
}
/*
- * Get the next entry in sequence from the BuildAccumulator's rbtree.
+ * Get the next entry in sequence from the BuildAccumulator's sorted hash
entries.
* This consists of a single key datum and a list (array) of one or more
* heap TIDs in which that key is found. The list is guaranteed sorted.
*/
@@ -268,25 +268,18 @@ ginGetBAEntry(BuildAccumulator *accum,
OffsetNumber *attnum, Datum *key, GinNullCategory
*category,
uint32 *n)
{
- GinEntryAccumulator *entry;
- ItemPointerData *list;
-
- entry = (GinEntryAccumulator *) rbt_iterate(&accum->tree_walk);
+ GinSortEntry *entry;
- if (entry == NULL)
+ if (accum->current_pos >= accum->num_entries)
return NULL; /* no more entries */
- *attnum = entry->attnum;
- *key = entry->key;
- *category = entry->category;
- list = entry->list;
- *n = entry->count;
-
- Assert(list != NULL && entry->count > 0);
+ entry = &accum->sorted_entries[accum->current_pos];
+ accum->current_pos++;
- if (entry->shouldSort && entry->count > 1)
- qsort(list, entry->count, sizeof(ItemPointerData),
- qsortCompareItemPointers);
+ *attnum = entry->hashkey.attnum;
+ *key = entry->hashkey.key;
+ *category = entry->hashkey.category;
+ *n = entry->numItems;
- return list;
+ return entry->items;
}
diff --git a/src/include/access/gin_private.h b/src/include/access/gin_private.h
index 3c5fd6ba817..df6b44796c6 100644
--- a/src/include/access/gin_private.h
+++ b/src/include/access/gin_private.h
@@ -18,7 +18,6 @@
#include "common/int.h"
#include "catalog/pg_am_d.h"
#include "fmgr.h"
-#include "lib/rbtree.h"
#include "nodes/tidbitmap.h"
#include "storage/bufmgr.h"
@@ -420,26 +419,15 @@ extern void ginadjustmembers(Oid opfamilyoid,
List *functions);
/* ginbulk.c */
-typedef struct GinEntryAccumulator
-{
- RBTNode rbtnode;
- Datum key;
- GinNullCategory category;
- OffsetNumber attnum;
- bool shouldSort;
- ItemPointerData *list;
- uint32 maxcount; /* allocated size of list[] */
- uint32 count; /* current number of list[]
entries */
-} GinEntryAccumulator;
typedef struct
{
- GinState *ginstate;
- Size allocatedMemory;
- GinEntryAccumulator *entryallocator;
- uint32 eas_used;
- RBTree *tree;
- RBTreeIterator tree_walk;
+ GinState * ginstate;
+ Size allocatedMemory;
+ struct ginbuild_hash * hash;
+ struct GinSortEntry * sorted_entries;
+ uint32 num_entries;
+ uint32 current_pos;
} BuildAccumulator;
extern void ginInitBA(BuildAccumulator *accum);
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 1040a65bc14..51f8b52b3a8 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1107,7 +1107,8 @@ GinBuildShared
GinBuildState
GinChkVal
GinEntries
-GinEntryAccumulator
+GinHashEntry
+GinHashKey
GinIndexStat
GinLeader
GinMetaPageData
@@ -1127,6 +1128,7 @@ GinScanKeyData
GinScanOpaque
GinScanOpaqueData
GinSegmentInfo
+GinSortEntry
GinState
GinStatsData
GinTernaryValue
--
2.50.1 (Apple Git-155)