apurtell opened a new pull request, #2624:
URL: https://github.com/apache/phoenix/pull/2624

   Design document: 
https://gist.github.com/apurtell/3af244c582d94ba83c61d5eca2c8b8c5
   
   ---
   
   ## Overview
   
   PHOENIX-7998 integrates native vector similarity search into Apache Phoenix, 
combining high-dimensional vector representations with Phoenix's relational 
query engine, global indexing infrastructure, and distributed coprocessor 
execution model. The feature encompasses fixed dimension vector types, scalar 
distance metrics with infix operators, server side exact top-K nearest neighbor 
evaluation with distributed merge sorting, Inverted File (IVF) approximate 
secondary indexing, BSON document vector extraction, single cell immutable 
storage support, cost based optimizer plan selection, adaptive probe expansion 
to mitigate candidate starvation under relational filtering, durable drift 
scorecarding with background rebuilds, and cross version compatibility 
considerations.
   
   The implementation is organized into 14 sequential commits against `master` 
(base `f590b9d9e0`). That base includes PHOENIX-7876, which reworked `EXPLAIN`: 
`ExplainPlanAttributes` is now built exclusively through its builder, plans 
disclose the optimizer's index selection rule and its rejected candidates, 
per-type server projections are emitted as `SERVER <TYPE> PROJECTION <n>` with 
per-expression detail lines, the client-side pipeline is carried as an ordered 
`clientSteps` list, and `EXPLAIN (FORMAT JSON)` renders the attribute tree 
directly. Vector plan diagnostics are integrated into that model throughout, so 
every vector disclosure described below is reachable from both the plan text 
and the JSON view. Each phase introduces a self-contained layer of 
functionality, progressing logically from type representations and scalar 
kernels through query compilation, index storage, write path coprocessor 
changes, execution plans, lifecycle tasks, and compatibility considerations. 
The sec
 tions below walk through every commit in merge order, detailing the 
architectural design, concrete engine changes, and verification strategy.
   
   ---
   
   ## Phase 1: Vector Data Type and SQL Grammar
   
   (19 files, +2782/−23)
   
   ### Architecture and Type Design
   
   Dense vector data is represented in Phoenix as fixed dimension typed columns 
through the logical types `VECTOR(FLOAT, <dim>)` and `VECTOR(DOUBLE, <dim>)`. 
Rather than creating a bespoke physical format, these types leverage Phoenix's 
existing packed array physical layout: raw, contiguous binary bytes with no 
per-element delimiters, offset arrays, or framing headers. Because each 
primitive element has a known static width (4 bytes for single precision 
floats, 8 bytes for double precision floats), the physical byte length equals 
the element width multiplied by the declared dimension. This enables 
dimensionality to be validated and derived directly from the binary payload.
   
   Decoupling logical vector identity from the low level byte codec allows the 
engine to enforce vector-specific validation, route distance functions, 
evaluate optimizer eligibility, and expose accurate JDBC metadata without 
requiring alterations to Phoenix's underlying storage layouts.
   
   ### Compiler and Schema Implementation
   
   The type system is realized through `PVectorFloat` and `PVectorDouble`, new 
`PDataType` subclasses registered in `PDataTypeFactory`. These classes embed 
dimension constraints into column metadata, enforce dimensional uniformity 
during mutation binding, and delegate serialization and deserialization to 
primitive float and double array codecs. In the grammar (`PhoenixSQL.g`), 
column definitions within `CREATE TABLE` and index key constraints support 
vector specifications via the `vector_component_type` production, which 
resolves element types dynamically through `PDataTypeFactory.typeForVector()`.
   
   Vector metadata is piped through the compiler via `ColumnDef`, which 
captures the dimension from the AST and injects it into `PColumnImpl` for 
catalog serialization. `ValueSchema` computes field widths directly from 
element widths and dimensions, ensuring row key serialization and schema byte 
length calculations remain consistent. During DML processing, `UpsertCompiler` 
inspects input arrays to ensure their lengths match the target column's 
declared dimension, throwing `SQLExceptionCode.VECTOR_DIMENSION_MISMATCH` if 
dimensions diverge. Across compilation and query optimization, 
`PArrayDataType.isVectorType()` provides a type predicate to distinguish vector 
types from general SQL arrays, while `SchemaUtil` propagates vector attributes 
during table and schema construction.
   
   ### Verification and Test Coverage
   
   Grammar parsing is verified by `VectorColumnParseTest` (19 tests), covering 
valid and invalid element types, missing or non-positive dimensions, and 
interactions with column constraints and modifiers. Codec integrity is tested 
in `VectorDataTypeTest` (43 tests), covering binary round-trips across element 
types, sort orders, nullability semantics, dimension validation, array 
coercion, and catalog serialization. End-to-end integration is validated by 
`VectorColumnIT` (12 tests) on a minicluster, confirming DDL creation, upsert 
binding, JDBC metadata inspection, null handling, and dimension mismatch 
rejection. Global regression suites (`PDataTypeTest`, `IndexUtilTest`) confirm 
type registration stability across Phoenix's global type dictionary.
   
   ---
   
   ## Phase 2: Distance Functions and Operators
   
   (23 files, +3038/−5)
   
   ### Geometric Metrics and Operator Semantics
   
   Similarity search relies on scalar distance functions that map vector pairs 
to `DOUBLE` precision distances. Following database similarity conventions, all 
metrics adhere to a nearest-first ordering rule where smaller output values 
denote greater similarity. The engine provides four foundational metrics: 
Euclidean distance (`L2_DISTANCE`) measures straight line spatial proximity, 
while squared Euclidean distance (`L2_DISTANCE_SQUARED`) computes the sum of 
squared coordinate differences, omitting the square root operation to minimize 
CPU cycles during nearest neighbor ranking. Directional divergence is evaluated 
using cosine distance (`COSINE_DISTANCE`), capturing angular separation 
independent of vector magnitude. For dot product similarity, inner product 
distance (`INNER_PRODUCT`) computes the negative dot product, negating the 
result so that standard ascending sort order aligns with maximum inner product 
similarity.
   
   To provide an ergonomic SQL interface, `PhoenixSQL.g` introduces infix 
operators: `<->` for Euclidean distance, `<=>` for cosine distance, and `<#>` 
for inner product. The lexer and parser normalize these operator tokens 
directly into function call parse trees via `ParseNodeFactory`.
   
   ### Computational Kernels and Hardware Acceleration
   
   Distance computations are optimized to eliminate object allocation. 
Functions process raw byte buffers directly without deserializing float or 
double elements onto the Java heap, allocating only the primitive `Double` 
result. The class hierarchy is anchored by `DistanceFunction`, an abstract base 
class that verifies dimension equality, unpackages byte representations, and 
delegates computation to concrete subclasses (`L2DistanceFunction`, 
`L2DistanceSquaredFunction`, `CosineDistanceFunction`, and 
`InnerProductDistanceFunction`). `DistanceFunction.isStateless()` returns 
`true`, qualifying distance evaluations for server-side coprocessor push-down. 
Each function is registered with an ordinal in `ExpressionType`.
   
   Vector computations are structured across a multi-release JAR layout:
   
   - **Java 11 Baseline:** Provided by `VectorDistanceUtil` under 
`src/main/java`, operating on raw `byte[]` and `float[]` buffers. This utility 
incorporates early-terminating variants (`*WithBound`) that abort accumulation 
as soon as the partial distance exceeds a known upper bound, avoiding unneeded 
arithmetic during bounded heap scans.
   
   - **Java 21 SIMD Overlay:** A multi-release overlay in `src/main/java21` 
provides a drop-in replacement `VectorDistanceUtil` that delegates to 
`PanamaDistanceKernel`. This kernel uses the JDK Vector API 
(`jdk.incubator.vector.FloatVector`) to execute 256-bit and 512-bit SIMD vector 
instructions for L2, dot-product, and cosine calculations while presenting an 
identical public API to the Java 11 class.
   
   The build configuration in `phoenix-core-client/pom.xml` introduces a 
`java21` Maven profile targeting `--release 21 --add-modules 
jdk.incubator.vector`, packaging versioned classes under 
`META-INF/versions/21`, while `pom.xml` passes incubator flags to Surefire for 
test execution.
   
   ### Verification and Test Coverage
   
   `DistanceFunctionsTest` (27 tests) validates end-to-end evaluation, 
dimensional compatibility checks, null semantics, and numerical accuracy 
against independently computed floating-point baselines for both `FLOAT` and 
`DOUBLE` vectors. Low-level scalar algorithms and early-termination bounds are 
verified by `VectorDistanceUtilTest` (3 tests). Operator tokenization, 
precedence relative to arithmetic and comparison operators, and AST 
normalization are tested in `DistanceOperatorParseTest` (20 tests). Type 
coercion between vectors and arrays is validated in `CoerceExpressionTest`.
   
   ---
   
   ## Phase 3: Exact Nearest-Neighbor Search Path
   
   (19 files, +2524/−51)
   
   ### Query Optimization and Plan Formulation
   
   Exact vector search provides brute force nearest-neighbor discovery when no 
vector index is available or when high selectivity from accompanying relational 
filters makes a direct table scan more efficient than index traversal. The 
query optimizer detects vector search patterns during compilation: 
`QueryCompiler` and `OrderByCompiler` analyze the query plan, identifying 
queries with an explicit `LIMIT`, an `ORDER BY` clause sorting ascending on a 
recognized distance function, a column reference pointing to a vector column, 
and a literal or bound parameter representing the constant query vector.
   
   When matched, `VectorSearchUtil` extracts a `VectorSearchDescriptor`. This 
descriptor captures the vector expression, query vector byte payload, 
dimension, distance metric (mapped between `DistanceFunctionType` and 
`DistanceMetric`), limit K, and source column references. Rather than 
transmitting all matching table rows to the client for sorting, the optimizer 
pushes the distance evaluation and bounded heap maintenance down to the 
RegionServers as scan attributes.
   
   ### Distributed Execution and Streaming Merge
   
   During scan execution, each RegionServer evaluates relational predicates, 
calculates distances using the pushed down expression, and tracks candidates 
within a local bounded heap of capacity K, returning only its local top-K rows. 
On the client, `MergeSortTopNResultIterator` coordinates a streaming K-way 
merge across all participating region scanners, guaranteeing global ordering 
without materializing full result sets in memory.
   
   Observability is integrated via `ExplainPlanAttributes` and `ExplainTable`. 
A recognized nearest-neighbor plan replaces the generic `SERVER TOP <n> ROWS 
SORTED BY [...]` step with `SERVER TOP-<K> BY <distance expression>`, and 
`MergeSortTopNResultIterator` replaces `CLIENT MERGE SORT` plus a separate 
`CLIENT LIMIT <K>` with the single step `CLIENT MERGE SORT TOP-<K>`, since the 
top-K bound is intrinsic to the merge rather than a post-merge truncation. Both 
steps set the `vectorSearch` flag and the `serverSortAlgo` attribute, and the 
client step is also appended to the ordered `clientSteps` pipeline, so the 
disclosure survives into `EXPLAIN (FORMAT JSON)`. At the AST layer, 
`DistanceFunctionParseNode` exposes metric metadata during semantic analysis. 
In addition, `PArrayDataType` adds automatic coercion from generic SQL arrays 
to vector types, allowing applications to pass vector literals or bind 
parameters as standard arrays without explicit casting.
   
   ### Verification and Test Coverage
   
   Descriptor extraction, query pattern recognition, constant folding, and 
error handling (such as missing limits, invalid sort directions, or mismatched 
dimensions) are tested by `VectorSearchUtilTest` (54 tests). Expression 
serialization is verified by `DistanceFunctionSerializationTest` (6 tests). 
Array-to-vector coercion rules are validated across types and dimensions by 
`PArrayDataTypeVectorCoercionTest` (23 tests). End-to-end cluster execution is 
proven by `VectorExactSearchIT` (13 tests), which exercises server-side 
push-down, cross-region streaming merges, all four distance metrics, salted and 
unsalted primary keys, bind parameters, null handling, query offsets, and 
explain plan verification on a live mini-cluster. Type round-trip assertions 
are reinforced in `PDataTypeTest`.
   
   ---
   
   ## Phase 4: System Catalog and Metadata Extensions
   
   (19 files, +1943/−29)
   
   ### Catalog Schema and Data Model
   
   Approximate vector indexing requires persistent metadata beyond that of 
standard secondary indexes: the distance metric, indexing algorithm, vector 
dimensionality, number of IVF posting lists, training sample size, and active 
centroid generation. This phase establishes the schema model, system catalog 
storage, and protobuf serialization protocols required to manage vector indexes.
   
   In the object model, `PTable` introduces `IndexType.VECTOR_GLOBAL` and 
exposes accessors for vector metadata: `vectorIndexAlgorithm`, 
`vectorDistanceMetric`, `vectorDimension`, `vectorIvfLists`, 
`vectorIvfSampleSize`, and `vectorCentroidGeneration`. `PTableImpl` and its 
builder handle the serialization, deserialization, and immutability contracts 
for these properties, while `DelegateTable` provides pass-through delegation. 
On the wire, `PTable.proto` defines protobuf optional fields 59 through 64 to 
transport vector index metadata between clients, query services, and 
RegionServers.
   
   ### System Tables and Lifecycle Management
   
   Vector metadata is stored across two system tables:
   1. `SYSTEM.CATALOG`: Extended with dedicated columns holding vector 
configuration parameters, managed by `MetaDataEndpointImpl` during table 
creation (`createTable`), schema lookups (`getTable`), and drops (`dropTable`). 
Constants for these columns are registered in `PhoenixDatabaseMetaData`.
   2. `SYSTEM.VECTOR_CENTROID`: A dedicated system table created to store 
trained centroid coordinates, generation summaries, and drift scorecards. As 
defined in `QueryConstants`, the table uses a composite primary key of 
`(INDEX_NAME, GENERATION_ID, CENTROID_ID)`, ensuring that centroids for a given 
generation are clustered together in storage.
   
   To accommodate these schema additions safely across upgrades, 
`MetaDataProtocol` advances `MIN_SYSTEM_TABLE_TIMESTAMP_5_4_0` from 
`MIN_TABLE_TIMESTAMP + 45` to `+ 51`. Both `ConnectionQueryServicesImpl` (for 
live clusters) and `ConnectionlessQueryServicesImpl` (for unit tests) bootstrap 
`SYSTEM.VECTOR_CENTROID` alongside existing system tables. Specific error 
scenarios during vector index creation are cataloged in `SQLExceptionCode`.
   
   ### Verification and Test Coverage
   
   `VectorSystemCatalogIT` (2 tests) verifies bootstrap creation and column 
layout of `SYSTEM.VECTOR_CENTROID`. Persistence, retrieval, sentinel row 
formatting, and multi-generation isolation are tested in 
`VectorCentroidTableIT` (20 tests). Catalog DDL round-trips and `PTable` 
accessor accuracy are verified in `VectorIndexIT` (442 lines). Enum 
serialization and index type predicates are covered by `VectorIndexTypeTest` 
(13 tests). Catalog metadata serialization is validated in 
`VectorDataTypeTest`. System table migration and bootstrapping regression 
suites (`MigrateSystemTablesToSystemNamespaceIT`, 
`SkipSystemTablesExistenceCheckIT`, and `SystemTablesCreationOnConnectionIT`) 
are updated and verified against the expanded system table set.
   
   ---
   
   ## Phase 5: Vector Index DDL and Grammar
   
   (12 files, +1514/−28)
   
   ### DDL Syntax and Language Grammar
   
   Vector indexes are created and modified using explicit SQL DDL syntax:
   
   ```sql
   CREATE VECTOR INDEX idx_name ON table_name (vector_expression)
   INCLUDE (covered_columns)
   WITH (metric='L2', algorithm='IVF', lists=1024, sample_size=50000)
   ```
   
   The grammar in `PhoenixSQL.g` introduces the `create_vector_index_node` 
production, adds support for `ALTER VECTOR INDEX`, and reserves `VECTOR` as a 
contextual keyword within DDL statements. The AST captures these specifications 
in `CreateIndexStatement`, which provides structured getters for the vector 
expression, metric, algorithm, list count, sample size, and parsed `WITH` 
properties.
   
   ### Compilation and Schema Validation
   
   Semantic validation is enforced by `CreateIndexCompiler`. The compiler 
verifies that:
   - Exactly one vector expression or column is declared in the index key.
   - The specified distance metric corresponds to a supported `DistanceMetric` 
enumeration.
   - The index algorithm is set to `IVF` (the supported approximate index 
algorithm).
   - Posting list counts and training sample sizes fall within valid numeric 
boundaries.
   - The target is a physical table; vector indexing over views is explicitly 
rejected.
   
   Once validated, `MetaDataClient` orchestrates index creation. It resolves 
the base table, confirms the presence and type of referenced columns, 
constructs the index schema by registering the vector expression as a 
functional index column, and invokes `MetaDataEndpointImpl` to write the vector 
metadata and `IndexType.VECTOR_GLOBAL` type marker into `SYSTEM.CATALOG`.
   
   ### Verification and Test Coverage
   
   DDL compilation rules are thoroughly exercised by `VectorIndexCompilerTest` 
(18 tests), confirming the acceptance of valid DDL, rejection of invalid 
metrics or unsupported algorithms, handling of missing parameters, dimension 
extraction from both physical columns and functional expressions, covered 
column inclusion, and `IF NOT EXISTS` idempotency. AST construction for complex 
vector expressions in index keys is validated by `VectorColumnParseTest` (15 
tests). End-to-end DDL execution against live clusters is tested in 
`VectorIndexIT` (432 lines), with additional serialization checks in 
`VectorIndexTypeTest` and `VectorDataTypeTest`.
   
   ---
   
   ## Phase 6: IVF Centroid Management
   
   (26 files, +8142/−6)
   
   ### Clustering Architecture and Metric Adaptation
   
   Inverted File (IVF) indexes partition high dimensional space into Voronoi 
cells using k-means clustering. Centroids define the centers of these cells, 
serving as partition keys for index rows. This phase implements the training 
pipeline, storage model, and caching tiers for centroid management across both 
local and distributed execution environments.
   
   Clustering is driven by `KMeansTrainer`, configured via `KMeansConfig`. The 
algorithm initializes cluster centers using k-means++ seeding, evaluates 
convergence across configurable iteration limits and tolerance thresholds, and 
tailors updates to the chosen distance metric:
   - **Euclidean (L2 / L2_SQUARED):** Standard Lloyd's algorithm computing the 
arithmetic mean of assigned vectors.
   - **Cosine:** Spherical k-means, computing the arithmetic mean and 
subsequently normalizing each centroid vector back to unit length after each 
iteration.
   - **Inner Product:** L2-based assignment using arithmetic mean updates.
   
   To counteract posting list imbalance caused by data clustering in real world 
workloads, `KMeansTrainer` performs post training rebalancing. Clusters 
exceeding variance thresholds are split, and empty clusters are reassigned, 
with the final effective centroid count and convergence statistics recorded in 
`KMeansResult`. Cluster health metrics (coefficient of variation, 
maximum-to-minimum size ratios, empty cluster counts) are computed and 
serialized by `ClusterSkewMetrics`.
   
   ### Distributed Training and Multi-Tier Caching
   
   For large tables where in-memory training is infeasible, clustering is 
distributed across MapReduce via `KMeansDistributedSampler`, 
`KMeansIterationDriver`, and `KMeansTool`. The distributed sampler runs 
reservoir sampling per mapper across base table region splits, while the 
iteration driver orchestrates distributed expectation-maximization iterations 
across the cluster. Vectors and centroids are serialized through Hadoop 
execution using `CentroidWritable` and `VectorWritable`, with parameters 
managed through `PhoenixConfigurationUtil`.
   
   Centroid persistence is governed by `CentroidManager`, which interfaces with 
`SYSTEM.VECTOR_CENTROID`. Centroids are written identified by the current 
generation, and a sentinal row holds per-generation training statistics. 
`CentroidManager` also handles lifecycle management, retiring older generations 
using efficient key-prefix range deletes.
   
   To achieve low query latency, `VectorCentroidCache` maintains an LRU cache 
of centroid arrays on both clients and RegionServers, keyed by `(indexName, 
generation)`. The cache avoids concurrent cache misses for the same generation. 
During DDL execution (`CREATE VECTOR INDEX`), `MetaDataClient` samples the base 
table and executes local training. If the base table is empty or contains fewer 
non-null vectors than the requested centroid count, `MetaDataClient` records 
the index metadata in `SYSTEM.CATALOG` while leaving the index in a building 
state with no centroid generation, deferring training to an asynchronous 
population build via `ServerBuildIndexCompiler`. Default cache sizes, local 
training limits, and brute force evaluation thresholds are configured in 
`QueryServicesOptions`.
   
   ### Verification and Test Coverage
   
   Algorithmic correctness, convergence guarantees, metric-specific spherical 
normalization, cluster splitting, and k-means++ reproducibility from fixed 
random seeds are validated by `KMeansTrainerTest` (15 tests). Centroid CRUD 
operations, sentinel row encoding, concurrent generation coexistence, and 
range-delete retirement are tested in `CentroidManagerTest` (18 tests). Cache 
behavior, including concurrency collapsing, LRU eviction, and generation 
isolation, is verified by `VectorCentroidCacheTest` (13 tests). Hadoop Writable 
serialization is covered by `CentroidWritableTest` (3 tests). End-to-end 
distributed MapReduce training on mini-clusters is tested in 
`KMeansDistributedTrainerIT` (8 tests). DDL-time training and persistence flows 
are validated in `VectorCentroidTableIT` (110 lines) and `VectorIndexIT` (99 
lines).
   
   ---
   
   ## Phase 7: Server-Side Write Pipeline for IVF
   
   (24 files, +2790/−234)
   
   ### Row Key Design and Posting List Layout
   
   The performance of IVF indexing depends on organizing index entries into 
contiguous physical posting lists. The index row key is structured to cluster 
vectors assigned to the same centroid:
   
   ```
   [Salt Byte (optional)] [Tenant ID (optional)] [Centroid ID] [Base Table 
Primary Key]
   ```
   
   Placing `Centroid ID` immediately after optional salt and tenant prefixes 
guarantees that all vectors assigned to a given centroid reside in a contiguous 
lexicographical range. Scanning an IVF partition translates directly into an 
efficient range scan over that centroid's posting list.
   
   ### Coprocessor Write Path and Mutation Lifecycle
   
   Index maintenance is coordinated across `IndexMaintainer`, coprocessors, and 
client paths. During table mutations, the vector expression is evaluated 
against the row data, its nearest centroid is identified using the active 
generation retrieved from `VectorCentroidCache`, and the index row key is 
constructed.
   
   On mutable tables, write path maintenance executes inside 
`IndexRegionObserver.preBatchMutate()`. The coprocessor evaluates both the 
before-image and after-image of the modified row:
   - If an insert occurs, it writes a new index row under the assigned centroid.
   - If an update alters the vector such that it moves across a Voronoi 
boundary into a different centroid, `IndexMaintainer` generates a paired 
mutation, a `Delete` for the old centroid's index key and a `Put` for the new 
centroid's index key.
   - If covered columns change but the centroid remains identical, an in-place 
update of the existing index row occurs.
   
   For immutable and transactional tables, this assignment is performed on the 
client. Serialization of vector metadata within index maintainers across RPC 
boundaries is handled by six new optional fields in 
`ServerCachingService.proto`. Read repair is integrated through 
`GlobalIndexChecker`, which validates that an index row's centroid assignment 
matches the vector value in the base table under the active generation, 
scheduling repairs for inconsistent entries.
   
   Bulk loading and index population are handled by `IndexTool` and 
`PhoenixIndexImportDirectMapper`, which configure centroid caches on mappers, 
project vector expressions from base table scans, and emit centroid keyed rows. 
Supporting updates include `DeleteCompiler`, `PostIndexDDLCompiler`, and 
`PhoenixIndexBuilder`. In `IndexUtil`, `findVectorColumn()` resolves vector 
expressions within index schemas, prioritizing functional expression columns 
before direct column references. In addition, `PVectorFloat` and 
`PVectorDouble` add sort-order-aware deserialization (`toObject(byte[], 
SortOrder)`), and `PDataTypeFactory` registers reverse mappings for JDBC 
metadata.
   
   ### Verification and Test Coverage
   
   `IndexMaintainerTest` (29 tests) covers vector index maintainer 
serialization, row key encoding across salted and multi-tenant tables, and 
centroid reassignment detection. End-to-end write behavior is tested in 
`VectorIndexWriteIT` (22 tests), verifying inserts, updates, deletes, centroid 
reassignment, covered column updates, partial row updates, null transitions, 
and `GlobalIndexChecker` read repairs on a live mini-cluster. Comprehensive 
write integration is reinforced in `VectorIndexIT` (951 lines). Vector column 
resolution ordering is tested in `IndexUtilVectorColumnTest` (6 tests). Type 
conversions and cache interactions under write load are validated in 
`ColumnInfoToStringEncoderDecoderTest`, `VectorIndexTypeTest`, 
`VectorDataTypeTest`, and `VectorCentroidCacheTest`.
   
   ---
   
   ## Phase 8: IVF Query Execution Path
   
   (37 files, +7126/−1966)
   
   ### Two-Phase Query Execution Architecture
   
   IVF similarity queries execute in a coordinated two-phase pipeline between 
the client and RegionServers, implemented in `VectorIndexScanPlan`:
   1. **Client Centroid Probing:** The client retrieves all centroids for the 
active generation from `VectorCentroidCache` and computes scalar distances 
against the query vector. It selects the P closest centroids, where P is the 
probe count (configured via session properties, table defaults, or 
`VECTOR_PROBE` hints in `HintNode`). `VectorSearchUtil` translates these 
centroids into multi-range skip-scan filters, expanding them across all salt 
buckets and aligning them with tenant prefixes.
   2. **Server-Side Bounded Scanning:** RegionServers receive the scan request 
via `NonAggregateRegionScannerFactory`. Instead of scanning the entire index, 
scanners read only the posting lists corresponding to the selected centroids.
   
   Within the RegionServer, `OrderedResultIterator` executes a two-phase 
scoring loop:
   - **Coarse Pass:** The scanner evaluates candidate vectors using early 
terminating bounds (`VectorDistanceUtil.*WithBound`). The bound is dynamically 
adjusted to the distance of the current K-th worst candidate in the local 
bounded heap. Vectors whose partial distance exceeds this threshold are 
discarded immediately, avoiding unnecessary floating point operations.
   - **Exact Pass:** Candidates that survive the coarse bound are scored 
completely, updating the local heap.
   
   RegionServers return their top candidates (scaled by an oversample factor 
configured via `ReadOnlyProps`) to the client, where `VectorIndexScanPlan` 
performs a distributed merge-sort to produce the final top-K result set.
   
   ### Optimizer Integration and Test Suite Modernization
   
   In `QueryOptimizer`, vector index eligibility is evaluated by comparing 
distance metrics, vector dimensions, and projected columns against the query's 
`VectorSearchDescriptor`. When an index matches, 
`IndexExpressionParseNodeRewriter` maps the base table vector expression to the 
index's functional column. `ExplainTable` renders execution details, including 
centroid counts, probe counts, distance metrics, oversample factors, and 
skip-scan ranges, and reports a vector index as `INDEX <name> VECTOR GLOBAL` 
alongside the existing `LOCAL`, `GLOBAL`, and `UNCOVERED GLOBAL` kinds. Index 
candidates that cannot serve the query are rejected with typed reasons drawn 
from `OptimizerReasons` (`not a nearest-neighbor search`, `distance metric does 
not match index`, `indexes a different vector column`, `vector dimension does 
not match index`, and `does not index the query vector expression`), which 
`EXPLAIN` discloses in `VERBOSE` mode as `/* !INDEX <name> -- <reason> */` 
lines rather than silen
 tly dropping the candidate. `ScanUtil` and `IndexTool` are updated to support 
vector scan attributes.
   
   ### Integration Test Refactor
   
   This phase also reorganizes the integration test suite, splitting large 
monolithic test files into specialized classes: `VectorExactSearchIT`, 
`VectorIndexWriteIT`, `VectorIvfLifecycleIT`, and `VectorIvfQueryIT`, supported 
by the shared fixture utility `VectorIndexTestUtil`.
   
   ### Verification and Test Coverage
   
   Query correctness, probe selection, skip-scan filtering, covered versus 
uncovered projections, multi-region streaming merges, double-precision vectors, 
oversample variations, and probe hints are validated by `VectorIvfQueryIT` (34 
tests), which includes recall assertions against brute force ground truth. 
Index lifecycle operations (creation, asynchronous builds, drops, rebuilds, and 
base table mutation interactions) are covered by `VectorIvfLifecycleIT` (14 
tests). Exact search pushdown and cross-region merges are verified in 
`VectorExactSearchIT` (13 tests). Unit tests in `VectorIndexScanPlanTest` (19 
tests) validate skip-scan range building, salt expansion, tenant alignment, and 
centroid ranking. The two phase bounded iterator is tested in 
`OrderedResultIteratorTest` (7 tests). Extended numerical accuracy and helper 
tests are added to `DistanceFunctionsTest`, `VectorSearchUtilTest`, 
`CentroidWritableTest`, and `VectorIndexWriteIT` (22 tests).
   
   ---
   
   ## Phase 9: Cost-Based Plan Selection and Explain Enhancements
   
   (10 files, +1233/−92)
   
   ### Cost Model Formulation
   
   To decide whether to execute an exact table scan or an approximate IVF index 
scan, `QueryOptimizer` incorporates a cost model implemented in `CostUtil`. The 
optimizer estimates execution costs by weighing five factors:
   1. **Candidate Cardinality:** Estimated using Phoenix table statistics and 
guideposts.
   2. **Relational Filter Selectivity:** The selectivity of accompanying 
`WHERE` clause predicates.
   3. **Posting List Scan Cost:** Proportional to the probe fraction (P / C), 
representing the fraction of total centroids C probed by the query.
   4. **Base Table Lookup Penalty:** For uncovered queries, the cost of 
performing random point lookups against the base table for non-indexed columns.
   5. **Covered Column Advantage:** The cost savings achieved when all 
requested columns are included directly within the index.
   
   ### Explain Plan Observability
   
   `VectorIndexScanPlan` overrides `getExplainPlan`, copying the attributes 
produced by its superclass into a fresh builder and adding the vector-specific 
disclosures on top. `ExplainTable` and the JSON renderer then expose:
   - Which relation was chosen, and why: the scanned table plus the `indexRule` 
selection label (`cost-based` when the cost model decided it) and, in `VERBOSE` 
mode, each rejected candidate with its reason.
   - The probe ratio, as a `CLIENT PROBING <P> OF <C> CENTROIDS (<metric>)` 
step and as the structured `vectorProbeCount`, `vectorCentroidCount`, and 
`vectorDistanceMetric` attributes.
   - Server-side two-phase rescoring, as `SERVER RESCORE TOP-<K> OF <n> 
CANDIDATES` and the `serverRescoreInfo` attribute, emitted only when the 
oversample factor exceeds 1.0.
   - The deferred base table join for an uncovered projection, as a `CLIENT 
MERGE [<cols>] FOR TOP-<K> ROWS` step spliced in directly after the merge sort, 
mirrored into `clientSteps` so the JSON pipeline keeps the same ordering, with 
the columns themselves in `clientMergeColumns`.
   
   ### Verification and Test Coverage
   
   Plan selection accuracy is verified in `VectorIvfCostIT` (4 tests), 
confirming that the optimizer selects IVF plans when covered indexes exist, 
falls back to exact scans when distance metrics diverge or index costs exceed 
table scans, and prioritizes covered indexes over uncovered alternatives. 
Cost-driven query execution is reinforced across 301 lines of tests in 
`VectorIvfQueryIT`, with cost calculation edge cases validated in 
`VectorSearchUtilTest`.
   
   ---
   
   ## Phase 10: BSON Vector Extraction
   
   (18 files, +1431/−63)
   
   ### Document Vector Extraction Semantics
   
   Modern architectures often store embeddings within JSON or BSON documents. 
Phoenix supports this pattern through `BSON_VECTOR_VALUE(document_column, 
'dotted.field.path', dimension)`, which extracts a typed `VECTOR(FLOAT, <dim>)` 
directly from a BSON document.
   
   The function expects binary data formatted according to BSON binary subtype 
9 using Float32 encoding. Implemented in `BsonVectorValueFunction` and parsed 
via `BsonVectorValueParseNode` (registered in `ExpressionType`), the function 
enforces strict data integrity rules:
   - If the dotted field path does not exist in the document, it returns SQL 
`NULL`.
   - If the path matches but has an incorrect BSON binary subtype, invalid 
floating-point encoding, mismatched byte length, or dimension divergence, it 
raises an immediate runtime error.
   
   ### Functional Indexing over Document Embeddings
   
   Extracted vectors integrate directly into Phoenix's secondary indexing 
engine via functional vector indexes:
   
   ```sql
   CREATE VECTOR INDEX idx ON tbl (BSON_VECTOR_VALUE(doc, 'embedding', 128)) 
INCLUDE (...)
   ```
   
   `MetaDataClient` validates that the source column is a valid BSON data type 
during DDL compilation.
   
   On the write path, `IndexMaintainer` (+165 lines) evaluates the extraction 
against both before-image and after-image document states, correctly handling 
partial BSON document updates. If the extracted vector shifts across a Voronoi 
boundary, it emits paired delete/insert mutations to keep posting lists 
synchronized. During query planning, `QueryOptimizer` and `VectorSearchUtil` 
match queries referencing `BSON_VECTOR_VALUE` against functional index 
definitions, while `IndexExpressionParseNodeRewriter` remaps the expression to 
the index table's functional column. Scan projection and execution are 
supported by `ProjectionCompiler`, `NonAggregateRegionScannerFactory`, and 
`BaseScannerRegionObserverConstants.BSON_VECTOR_VALUE_FUNCTION`, with wire 
metadata carried in `ServerCachingService.proto`. `ProjectionCompiler` 
registers the pushed down extractions on the statement context, so they join 
`BSON_VALUE` in the shared `SERVER BSON PROJECTION <n>` clause with one 
indented detail line pe
 r server evaluated expression.
   
   ### Verification and Test Coverage
   
   Function evaluation, missing path handling, subtype validation, dimension 
mismatch enforcement, and expression serialization are tested by 
`BsonVectorValueFunctionTest` (16 tests). Parsing of BSON vector expressions in 
DDL statements is verified in `VectorColumnParseTest`. End-to-end integration 
is proven in `VectorIndexIT` (474 lines), validating index creation, bulk 
population, query execution, covered column retrieval, and document 
update-driven index maintenance.
   
   ---
   
   ## Phase 11: Single-Cell Storage Support for Vector Indexes
   
   (7 files, +731/−75)
   
   ### Storage Layout Compatibility
   
   Phoenix supports two storage layouts for immutable tables: standard 
`ONE_CELL_PER_COLUMN` and `SINGLE_CELL_ARRAY_WITH_OFFSETS`. In the single cell 
layout, all column values within a column family are packed into a single HBase 
cell alongside an offset array. This format optimizes disk footprints and read 
performance for dense wide tables. Vector indexes must function identically 
across both storage models.
   
   ### Maintainer Adaptations and Transcoding
   
   In `IndexMaintainer`, column expression evaluation was refactored to 
recognize and unwrap `SingleCellColumnExpression` wrappers. This enables the 
write pipeline to extract vector values from packed single cell cells during 
mutation processing. During partial updates of covered columns, the maintainer 
preserves vector values across row reconstruction phases. Furthermore, 
sort-order transcoding is normalized when reading vector columns from single 
cell tables, ensuring consistent binary keys regardless of physical storage 
format.
   
   Metadata propagation in `MetaDataClient` and `MetaDataEndpointImpl` was 
updated to ensure that immutable storage scheme attributes are preserved during 
vector index creation and administrative operations.
   
   ### Verification and Test Coverage
   
   Verification in `VectorIndexWriteIT` (271 lines) confirms that insert, 
update, delete, and centroid reassignment mutations generate identical index 
row keys under both `ONE_CELL_PER_COLUMN` and `SINGLE_CELL_ARRAY_WITH_OFFSETS`. 
Query consistency across storage schemes is proven in `VectorIvfQueryIT` (139 
lines), verifying that IVF search recall is unaffected by the underlying table 
format. Unit tests in `IndexMaintainerTest` (215 lines) validate expression 
unwrapping and row key generation, while DDL persistence is verified in 
`VectorIndexIT` (70 lines).
   
   ---
   
   ## Phase 12: Adaptive Probing and Hybrid Filtering
   
   (15 files, +1675/−44)
   
   ### Adaptive Probe Expansion
   
   When similarity queries include selective relational predicates, scanning 
only the initial P centroids can result in candidate starvation, a state where 
relational filters discard most candidates, returning fewer than the requested 
K results.
   
   Adaptive probing resolves this within `VectorIndexScanPlan`. The query 
begins by scanning posting lists for the initial P closest centroids. If 
post-filtering leaves fewer than K candidates, the plan initiates dynamic probe 
expansion:
   1. It computes distances to the remaining centroids to identify the next 
closest batch.
   2. It constructs new multi-range skip-scan filters for these posting lists.
   3. It dispatches supplementary scans to the RegionServers.
   4. It merges newly discovered candidates into the existing candidate pool 
using `MergeSortTopNResultIterator`.
   
   This cycle repeats until K candidates are accumulated, all viable centroids 
are scanned, or a configured maximum probe limit is reached. The maximum probe 
threshold can be specified per-query using the `VECTOR_MAX_PROBE` hint in 
`HintNode`, or globally via `QueryServices.VECTOR_MAX_PROBE_LIMIT` and 
`QueryServicesOptions`.
   
   ### Hybrid Filter-First Plan Selection
   
   In scenarios where a query includes a highly selective scalar predicate 
backed by an existing standard secondary index, scanning an IVF index even with 
adaptive probing can be suboptimal. For these workloads, `QueryOptimizer` 
introduces `FilterFirstIndexPlan` (supported by `ScanPlan` and 
`DelegateQueryPlan`).
   
   The filter-first plan executes in reverse order:
   1. It uses the scalar secondary index to locate rows that satisfy the 
selective predicate.
   2. It retrieves the vector columns for only those matching rows.
   3. It computes exact vector distances and ranks the candidates.
   
   The optimizer compares the estimated cost of `FilterFirstIndexPlan` against 
`VectorIndexScanPlan`, leveraging `VectorSearchUtil` for filter selectivity 
estimation and `BsonValueFunction` for document predicate analysis.
   
   ### Verification and Test Coverage
   
   Adaptive probe expansion, candidate starvation recovery, maximum probe 
bounds, filter-first plan selection, and covered/uncovered column interactions 
are verified in `VectorIndexIT` (736 lines). Unit tests in 
`VectorIndexScanPlanTest` (95 lines) validate batch range generation, 
cumulative range merging, and limit enforcement. Predicate extraction for 
filter-first plans is verified in `BsonValueFunctionTest` (95 lines), and 
selectivity estimation routines are tested in `VectorSearchUtilTest` (119 
lines).
   
   ---
   
   ## Phase 13: Drift Scorecarding and Background Rebuild
   
   (35 files, +6138/−144)
   
   ### Continuous Health Monitoring and Scorecard Storage
   
   Over time, shifts in vector data distributions cause cluster imbalance. Some 
posting lists grow disproportionately large while others become sparse, 
degrading scan efficiency and reducing search recall. Phoenix addresses this 
with an inline drift scorecarding and automated rebuild framework.
   
   Index health is tracked in `SYSTEM.VECTOR_CENTROID` as durable scorecard 
rows stored alongside the centroids they monitor (one row per centroid per 
generation, plus a sentinel summary row). The write path updates the scorecard 
inline. Because `IndexMaintainer` already inspects before-images and 
after-images to derive centroid assignments, it calculates posting list 
population deltas with negligible overhead.
   
   To prevent write amplification on the system catalog, RegionServers buffer 
deltas in memory using `ScorecardAccumulator`. Inside `IndexRegionObserver`, 
accumulators aggregate mutation counts across configurable time windows 
(`QueryServices.VECTOR_SCORECARD_FLUSH_INTERVAL_MS`), periodically flushing 
deltas to `SYSTEM.VECTOR_CENTROID` using atomic aggregating upserts 
(`ScorecardRow`).
   
   ### Counter Reconciliation and Drift Assessment
   
   Because node crashes or aborted transactions can introduce minor 
discrepancies in in-memory counters, Phoenix provides scheduled reconciliation. 
`VectorScorecardReconcileTask` (`TaskType.VECTOR_SCORECARD_RECONCILE` in 
`PTable`), registered in `TaskRegionObserver`, executes a lightweight grouped 
count scan along the index's leading key column. This scan recalculates exact 
posting list counts, correcting counter drift. Reconciliation runs on a 
background schedule (`QueryServices.VECTOR_SCORECARD_RECONCILE_INTERVAL_MS`) 
and is executed before any rebuild decision.
   
   Index drift is evaluated by `VectorIndexScorecard` against four 
administrative thresholds:
   1. **Posting List Skew Ratio:** Ratio of the largest posting list to the 
median list size.
   2. **Coefficient of Variation (CV):** Standard deviation of cluster sizes 
divided by the mean.
   3. **Empty Centroid Fraction:** Percentage of centroids with zero assigned 
vectors.
   4. **Reassignment Rate:** Rate at which document updates shift vectors 
across centroid boundaries.
   
   Assessment verdicts and diagnostic summaries are recorded in 
`GenerationSummary` within the sentinel row of `SYSTEM.VECTOR_CENTROID`.
   
   ### Zero-Downtime Rebuild and Migration
   
   When drift thresholds are crossed, or upon manual invocation via 
`MetaDataClient`, `VectorIndexRebuildTask` (`TaskType.VECTOR_INDEX_REBUILD`) 
coordinates an asynchronous background rebuild:
   1. **Model Training:** It samples the base table and trains a new centroid 
generation (G_new).
   2. **Data Migration:** Using `CentroidManager`, it scans the index and 
rewrites entries under the new centroid model using atomic paired delete/insert 
mutations. During migration, queries continue executing against the active 
generation (G_old).
   3. **Generation Promotion:** Once migration is complete, the task updates 
the index's active generation to G_new in `SYSTEM.CATALOG`, initializes the new 
generation's scorecard, and retires G_old.
   4. **Garbage Collection:** Obsolete centroid rows and scorecard data for 
G_old are removed via key-prefix range deletes.
   
   During rebuild operations, query probe counts can be dynamically adjusted 
according to configured rebuild probe policies 
(`QueryServices.VECTOR_REBUILD_PROBE_POLICY` and factor). `EXPLAIN` discloses 
that adjustment so a plan captured mid-rebuild is not mistaken for steady 
state: the `EXPAND` policy annotates the probe step as `CLIENT PROBING <P> OF 
<C> CENTROIDS (<metric>) (REBUILD IN PROGRESS: EXPANDED FROM <P_base>)`, and 
the `EXACT` policy replaces it with `CLIENT EXACT VECTOR EVALUATION OVER <C> 
CENTROIDS (REBUILD IN PROGRESS)`. Constants and DDL schemas for scorecard 
columns are registered in `PhoenixDatabaseMetaData` and `QueryConstants`.
   
   ### Verification and Test Coverage
   
   Scorecard evaluation logic, CV calculations, empty-cluster detection, and 
verdict generation are verified in `VectorIndexScorecardTest` (11 tests). 
Reconciliation accuracy against deliberately corrupted counters and task 
scheduling are tested in `VectorScorecardReconcileTaskTest` (5 tests). Centroid 
migration, generation promotion, and range retirement are validated in 
`CentroidManagerTest` (298 lines). Migration correctness and generation 
coexistence during live rebuilds are proven in `VectorIvfLifecycleIT` (362 
lines). Scorecard persistence and reconciliation integration are verified in 
`VectorCentroidTableIT` (589 lines). End-to-end drift detection, delta 
accumulation, flush verification, and automated background rebuilds on live 
clusters are tested in `VectorIndexIT` (1384 lines). Additional probe policy 
and scorecard assertions are covered in `VectorIndexScanPlanTest`, 
`VectorColumnIT`, and `VectorIndexWriteIT`.
   
   ---
   
   ## Phase 14: Compatibility and Version Validation
   
   (7 files, +252/−22)
   
   ### Cross-Version Compatibility and Protocol Gating
   
   In distributed environments running rolling upgrades or mixed client 
deployments, older clients that lack vector index awareness must be prevented 
from executing invalid plans or corrupting index metadata. Conversely, newer 
clients must handle connections to older server clusters cleanly.
   
   This phase implements strict version gating:
   - **Client-Side Validation:** `ConnectionQueryServices` and 
`ConnectionQueryServicesImpl` inspect cluster capabilities during connection 
initialization. Clients reject vector index metadata if the server reports a 
version below the vector support threshold.
   - **DDL Gating:** In `MetaDataClient`, attempts to create a vector index 
against a cluster below the minimum supported version fail immediately with an 
explicit compatibility error.
   - **Server-Side Protection:** `IndexRegionObserver` inspects index 
maintainers during mutation processing. If a maintainer contains vector 
specifications on a RegionServer version that predates vector support, it logs 
a warning and skips vector-specific handling rather than failing unexpectedly.
   
   ### Wire Protocol Backward Compatibility
   
   Protocol buffer serialization in `IndexMaintainer` (+51 lines) adheres to 
strict backward-compatibility rules. When serializing metadata for older peers, 
vector-specific fields are omitted to prevent deserialization errors. During 
deserialization, missing optional vector fields are interpreted as standard 
non-vector secondary indexes, ensuring unupgraded nodes continue normal 
operation without interruption.
   
   ### Verification and Test Coverage
   
   Version validation behavior is tested in `VectorIndexIT` (89 lines), 
verifying that older client simulations reject vector metadata, older server 
simulations fail vector DDL with descriptive errors, and protobuf serialization 
round-trips maintain wire compatibility. Protocol compatibility and 
version-gated serialization are validated in `VectorIndexTypeTest` (67 lines).
   
   ---
   
   ## Test Verification and Validation
   
   All test suites were executed against the branch head (`582ecfaece`) on 
macOS (Darwin 25.5.0, Apple Silicon) using JDK 11 (Temurin 11.0.30) for 
baseline verification and JDK 21 (Temurin 21.0.12.1) for Panama SIMD 
validation. The test counts and pass/fail outcomes below are from a 
re-execution after the branch was rebased onto base `f590b9d9e0`. The 
integration timings are from that run too, measured serially with one `mvn 
verify` per class. The unit timings are indicative only: the unit suite runs 8 
reused forks in parallel, so a class that lands in an already warm JVM reports 
a fraction of the time the same class takes in a cold one. The JDK 21 Panama 
overlay figures in the final subsection carry over from the pre-rebase 
verification, as the rebase touched no distance kernel.
   
   Because the vector attributes are additive to `ExplainPlanAttributes`, they 
are also exercised by the `EXPLAIN` regression suites that ship with the base 
commit. `ExplainPlanTest` compares a complete normalized JSON attribute tree 
for each of its 111 corpus cases, so every vector attribute is asserted, in its 
unset form, across all of them; `QueryOptimizerTest`, `QueryCompilerTest`, 
`ExplainOptionsParserTest`, and `ExplainJsonOutputTest` cover the index 
selection rule and rejected-candidate plumbing that vector index rejections 
feed into. Those five suites total 397 tests and pass, with 3 pre-existing 
skips.
   
   ### Unit Test Execution
   
   The unit test suite completed with **419 passed tests, 0 failures, and 0 
errors**, executed with the project default of 8 parallel forks 
(`numForkedUT`). One unrelated class, `PhoenixStatsCacheLoaderTest`, can stall 
under that level of fork contention: `testStatsBeingAutomaticallyRefreshed` 
waits on an untimed `CountDownLatch` for a timing-dependent stats cache 
refresh. It is pre-existing, touches no vector or `EXPLAIN` code, and passes in 
under a second when its class is run alone.
   
   | Test Class | Tests | Execution Time | Primary Test Coverage |
   |---|---|---|---|
   | `VectorSearchUtilTest` | 54 | 5.4s | Plan descriptor extraction, metric 
resolution, constant folding, edge cases |
   | `VectorDataTypeTest` | 43 | 0.2s | Codec round-trips, sort orders, 
nullability, dimension validation |
   | `PDataTypeTest` | 38 | 0.5s | Type registration, global type-iteration 
stability |
   | `IndexMaintainerTest` | 29 | 3.4s | Row key encoding, centroid 
reassignments, single-cell unwrapping |
   | `DistanceFunctionsTest` | 27 | 2.6s | End-to-end evaluation, dimensional 
compatibility, accuracy against baselines |
   | `PArrayDataTypeVectorCoercionTest` | 23 | 0.01s | Array-to-vector coercion 
rules across element types and dimensions |
   | `DistanceOperatorParseTest` | 20 | 1.5s | Infix operator parsing, 
precedence, AST normalization |
   | `VectorColumnParseTest` | 19 | 1.9s | Column definitions, DDL grammar, 
constraint parsing |
   | `VectorIndexScanPlanTest` | 19 | 2.9s | Skip-scan range building, salt 
expansion, tenant alignment, ranking |
   | `CentroidManagerTest` | 18 | 1.5s | Generation CRUD, sentinel rows, 
lifecycle retirement by range delete |
   | `VectorIndexCompilerTest` | 18 | 3.2s | DDL compilation, metric 
validation, algorithm validation, covered columns |
   | `BsonVectorValueFunctionTest` | 16 | 0.2s | BSON path extraction, binary 
subtype 9 validation, dimension checks |
   | `KMeansTrainerTest` | 15 | 0.7s | Convergence detection, cluster 
splitting, spherical k-means, seeding |
   | `VectorCentroidCacheTest` | 13 | 2.2s | Hit/miss ratios, LRU eviction, 
concurrency collapsing, generation isolation |
   | `VectorIndexTypeTest` | 13 | 0.1s | IndexType serialization, ordinal 
stability, version-gated serde |
   | `VectorIndexScorecardTest` | 11 | 2.2s | Scorecard evaluation, skew 
ratios, CV computation, reassignment rates |
   | `OrderedResultIteratorTest` | 7 | 2.5s | Two-phase scoring, 
early-termination bounding, heap limits |
   | `DistanceFunctionSerializationTest` | 6 | 0.03s | Serde round-trips 
through Phoenix expression serialization framework |
   | `IndexUtilVectorColumnTest` | 6 | 1.2s | Vector column resolution ordering 
(functional expressions vs. direct refs) |
   | `CoerceExpressionTest` | 5 | 0.06s | Vector-to-vector and array coercion 
expression rules |
   | `VectorScorecardReconcileTaskTest` | 5 | 0.9s | Counter reconciliation 
accuracy under corruption, task scheduling |
   | `VectorDistanceUtilTest` | 3 | 0.01s | Low-level scalar kernels (L2, 
cosine, dot product) on packed buffers |
   | `BsonValueFunctionTest` | 3 | 2.5s | BSON predicate extraction for 
filter-first plan selection |
   | `CentroidWritableTest` | 3 | 0.01s | Hadoop Writable serialization 
round-trips for centroid data |
   | `ColumnInfoToStringEncoderDecoderTest` | 3 | 0.1s | ColumnInfo 
serialization with vector data types |
   | `IndexUtilTest` | 2 | 1.4s | Secondary index utility constants and column 
lookup validation |
   
   ### Integration Test Execution
   
   The integration test suite completed with **214 passed tests, 0 failures, 0 
errors, and 1 skipped**, run strictly serially with one `mvn verify` invocation 
per class (total wall time approximately 43 minutes):
   
   | Test Class | Tests | Execution Time | Primary Test Coverage |
   |---|---|---|---|
   | `VectorIndexIT` | 63 | 6m 42s | End-to-end DDL, writes, queries, 
scorecarding, and background rebuilds |
   | `VectorIvfQueryIT` | 34 | 53s | IVF execution, multi-region merge, covered 
projections, recall verification |
   | `VectorIndexWriteIT` | 22 | 1m 39s | Write pipeline, centroid 
reassignment, single-cell storage, read repairs |
   | `VectorCentroidTableIT` | 20 | 5s | System table CRUD, generation 
isolation, sentinel rows, scorecard persistence |
   | `SystemTablesCreationOnConnectionIT` | 15 (1 skipped) | 11m 20s | System 
table bootstrap validation (includes `SYSTEM.VECTOR_CENTROID`) |
   | `VectorIvfLifecycleIT` | 14 | 1m 34s | Index build, async population, 
drop, rebuild, generation coexistence |
   | `VectorExactSearchIT` | 13 | 0.2s | Server push-down, cross-region 
streaming merge, offsets, bind variables |
   | `VectorColumnIT` | 12 | 25s | Vector column DDL, upsert binding, JDBC 
metadata, constraint enforcement |
   | `KMeansDistributedTrainerIT` | 8 | 1m 9s | Distributed MapReduce k-means 
training on mini-cluster |
   | `VectorIvfCostIT` | 4 | 39s | Cost-based plan selection (IVF vs. exact 
scan vs. covered index preference) |
   | `MigrateSystemTablesToSystemNamespaceIT` | 4 | 5m 57s | Namespace 
migration containing new system tables |
   | `SkipSystemTablesExistenceCheckIT` | 3 | 1m 10s | Fast connection 
bootstrap without redundant system table checks |
   | `VectorSystemCatalogIT` | 2 | 0.2s | Bootstrap creation and column 
verification of `SYSTEM.VECTOR_CENTROID` |
   
   The single skipped test corresponds to a pre-existing `@Ignore` annotation 
in `SystemTablesCreationOnConnectionIT`, which is unrelated to vector indexing.
   
   ### JDK 21 SIMD Overlay Verification
   
   The `java21` Maven profile introduces a multi-release JAR overlay with 
Panama Vector API distance kernels. Because Phoenix's default compile target is 
JDK 11, this overlay is maintained in `src/main/java21` and validated through a 
one-off verification procedure:
   1. **Compilation:** The overlay compiles cleanly under JDK 21 with 
`--release 21 --add-modules jdk.incubator.vector`.
   2. **API Parity:** The Java 21 `VectorDistanceUtil` implementation matches 
the public method signatures of the Java 11 base class, adhering to the 
multi-release JAR specification.
   3. **Algorithmic Equivalence:** 104 tests across five suites 
(`DistanceFunctionsTest`, `VectorDataTypeTest`, `BsonVectorValueFunctionTest`, 
`KMeansTrainerTest`, and `VectorDistanceUtilTest`) were re-run under JDK 21 
with the Panama overlay active on the classpath. All tests passed, confirming 
that the SIMD-accelerated kernels produce results identical to the scalar 
reference paths.
   
   ### Testing Notes
   
   The test harness incorporates several architectural testing standards:
   - **Deterministic Geometric Fixtures:** Rather than relying on 
non-deterministic random data, integration tests utilize geometric fixtures 
with mathematically predetermined centroid placements and known cluster 
memberships (`VectorIndexTestUtil`). This allows exact recall and Voronoi 
boundary transitions to be verified deterministically.
   - **Independent Ground-Truth Comparison:** Exact nearest-neighbor searches 
are validated against independently computed brute-force rankings to prevent 
circular verification.
   - **Exhaustive Code Coverage:** Across 39 modified and new test files 
containing 622 `@Test` annotations, all new execution paths (including failure 
modes, boundary transitions, and error codes) maintain full coverage without 
placeholder assertions or ignored test methods.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to