apurtell opened a new pull request, #2624: URL: https://github.com/apache/phoenix/pull/2624
Design document: https://gist.github.com/apurtell/3af244c582d94ba83c61d5eca2c8b8c5 --- ## Overview PHOENIX-7998 integrates native vector similarity search into Apache Phoenix, combining high-dimensional vector representations with Phoenix's relational query engine, global indexing infrastructure, and distributed coprocessor execution model. The feature encompasses fixed dimension vector types, scalar distance metrics with infix operators, server side exact top-K nearest neighbor evaluation with distributed merge sorting, Inverted File (IVF) approximate secondary indexing, BSON document vector extraction, single cell immutable storage support, cost based optimizer plan selection, adaptive probe expansion to mitigate candidate starvation under relational filtering, durable drift scorecarding with background rebuilds, and cross version compatibility considerations. The implementation is organized into 14 sequential commits against `master` (base `f590b9d9e0`). That base includes PHOENIX-7876, which reworked `EXPLAIN`: `ExplainPlanAttributes` is now built exclusively through its builder, plans disclose the optimizer's index selection rule and its rejected candidates, per-type server projections are emitted as `SERVER <TYPE> PROJECTION <n>` with per-expression detail lines, the client-side pipeline is carried as an ordered `clientSteps` list, and `EXPLAIN (FORMAT JSON)` renders the attribute tree directly. Vector plan diagnostics are integrated into that model throughout, so every vector disclosure described below is reachable from both the plan text and the JSON view. Each phase introduces a self-contained layer of functionality, progressing logically from type representations and scalar kernels through query compilation, index storage, write path coprocessor changes, execution plans, lifecycle tasks, and compatibility considerations. The sec tions below walk through every commit in merge order, detailing the architectural design, concrete engine changes, and verification strategy. --- ## Phase 1: Vector Data Type and SQL Grammar (19 files, +2782/−23) ### Architecture and Type Design Dense vector data is represented in Phoenix as fixed dimension typed columns through the logical types `VECTOR(FLOAT, <dim>)` and `VECTOR(DOUBLE, <dim>)`. Rather than creating a bespoke physical format, these types leverage Phoenix's existing packed array physical layout: raw, contiguous binary bytes with no per-element delimiters, offset arrays, or framing headers. Because each primitive element has a known static width (4 bytes for single precision floats, 8 bytes for double precision floats), the physical byte length equals the element width multiplied by the declared dimension. This enables dimensionality to be validated and derived directly from the binary payload. Decoupling logical vector identity from the low level byte codec allows the engine to enforce vector-specific validation, route distance functions, evaluate optimizer eligibility, and expose accurate JDBC metadata without requiring alterations to Phoenix's underlying storage layouts. ### Compiler and Schema Implementation The type system is realized through `PVectorFloat` and `PVectorDouble`, new `PDataType` subclasses registered in `PDataTypeFactory`. These classes embed dimension constraints into column metadata, enforce dimensional uniformity during mutation binding, and delegate serialization and deserialization to primitive float and double array codecs. In the grammar (`PhoenixSQL.g`), column definitions within `CREATE TABLE` and index key constraints support vector specifications via the `vector_component_type` production, which resolves element types dynamically through `PDataTypeFactory.typeForVector()`. Vector metadata is piped through the compiler via `ColumnDef`, which captures the dimension from the AST and injects it into `PColumnImpl` for catalog serialization. `ValueSchema` computes field widths directly from element widths and dimensions, ensuring row key serialization and schema byte length calculations remain consistent. During DML processing, `UpsertCompiler` inspects input arrays to ensure their lengths match the target column's declared dimension, throwing `SQLExceptionCode.VECTOR_DIMENSION_MISMATCH` if dimensions diverge. Across compilation and query optimization, `PArrayDataType.isVectorType()` provides a type predicate to distinguish vector types from general SQL arrays, while `SchemaUtil` propagates vector attributes during table and schema construction. ### Verification and Test Coverage Grammar parsing is verified by `VectorColumnParseTest` (19 tests), covering valid and invalid element types, missing or non-positive dimensions, and interactions with column constraints and modifiers. Codec integrity is tested in `VectorDataTypeTest` (43 tests), covering binary round-trips across element types, sort orders, nullability semantics, dimension validation, array coercion, and catalog serialization. End-to-end integration is validated by `VectorColumnIT` (12 tests) on a minicluster, confirming DDL creation, upsert binding, JDBC metadata inspection, null handling, and dimension mismatch rejection. Global regression suites (`PDataTypeTest`, `IndexUtilTest`) confirm type registration stability across Phoenix's global type dictionary. --- ## Phase 2: Distance Functions and Operators (23 files, +3038/−5) ### Geometric Metrics and Operator Semantics Similarity search relies on scalar distance functions that map vector pairs to `DOUBLE` precision distances. Following database similarity conventions, all metrics adhere to a nearest-first ordering rule where smaller output values denote greater similarity. The engine provides four foundational metrics: Euclidean distance (`L2_DISTANCE`) measures straight line spatial proximity, while squared Euclidean distance (`L2_DISTANCE_SQUARED`) computes the sum of squared coordinate differences, omitting the square root operation to minimize CPU cycles during nearest neighbor ranking. Directional divergence is evaluated using cosine distance (`COSINE_DISTANCE`), capturing angular separation independent of vector magnitude. For dot product similarity, inner product distance (`INNER_PRODUCT`) computes the negative dot product, negating the result so that standard ascending sort order aligns with maximum inner product similarity. To provide an ergonomic SQL interface, `PhoenixSQL.g` introduces infix operators: `<->` for Euclidean distance, `<=>` for cosine distance, and `<#>` for inner product. The lexer and parser normalize these operator tokens directly into function call parse trees via `ParseNodeFactory`. ### Computational Kernels and Hardware Acceleration Distance computations are optimized to eliminate object allocation. Functions process raw byte buffers directly without deserializing float or double elements onto the Java heap, allocating only the primitive `Double` result. The class hierarchy is anchored by `DistanceFunction`, an abstract base class that verifies dimension equality, unpackages byte representations, and delegates computation to concrete subclasses (`L2DistanceFunction`, `L2DistanceSquaredFunction`, `CosineDistanceFunction`, and `InnerProductDistanceFunction`). `DistanceFunction.isStateless()` returns `true`, qualifying distance evaluations for server-side coprocessor push-down. Each function is registered with an ordinal in `ExpressionType`. Vector computations are structured across a multi-release JAR layout: - **Java 11 Baseline:** Provided by `VectorDistanceUtil` under `src/main/java`, operating on raw `byte[]` and `float[]` buffers. This utility incorporates early-terminating variants (`*WithBound`) that abort accumulation as soon as the partial distance exceeds a known upper bound, avoiding unneeded arithmetic during bounded heap scans. - **Java 21 SIMD Overlay:** A multi-release overlay in `src/main/java21` provides a drop-in replacement `VectorDistanceUtil` that delegates to `PanamaDistanceKernel`. This kernel uses the JDK Vector API (`jdk.incubator.vector.FloatVector`) to execute 256-bit and 512-bit SIMD vector instructions for L2, dot-product, and cosine calculations while presenting an identical public API to the Java 11 class. The build configuration in `phoenix-core-client/pom.xml` introduces a `java21` Maven profile targeting `--release 21 --add-modules jdk.incubator.vector`, packaging versioned classes under `META-INF/versions/21`, while `pom.xml` passes incubator flags to Surefire for test execution. ### Verification and Test Coverage `DistanceFunctionsTest` (27 tests) validates end-to-end evaluation, dimensional compatibility checks, null semantics, and numerical accuracy against independently computed floating-point baselines for both `FLOAT` and `DOUBLE` vectors. Low-level scalar algorithms and early-termination bounds are verified by `VectorDistanceUtilTest` (3 tests). Operator tokenization, precedence relative to arithmetic and comparison operators, and AST normalization are tested in `DistanceOperatorParseTest` (20 tests). Type coercion between vectors and arrays is validated in `CoerceExpressionTest`. --- ## Phase 3: Exact Nearest-Neighbor Search Path (19 files, +2524/−51) ### Query Optimization and Plan Formulation Exact vector search provides brute force nearest-neighbor discovery when no vector index is available or when high selectivity from accompanying relational filters makes a direct table scan more efficient than index traversal. The query optimizer detects vector search patterns during compilation: `QueryCompiler` and `OrderByCompiler` analyze the query plan, identifying queries with an explicit `LIMIT`, an `ORDER BY` clause sorting ascending on a recognized distance function, a column reference pointing to a vector column, and a literal or bound parameter representing the constant query vector. When matched, `VectorSearchUtil` extracts a `VectorSearchDescriptor`. This descriptor captures the vector expression, query vector byte payload, dimension, distance metric (mapped between `DistanceFunctionType` and `DistanceMetric`), limit K, and source column references. Rather than transmitting all matching table rows to the client for sorting, the optimizer pushes the distance evaluation and bounded heap maintenance down to the RegionServers as scan attributes. ### Distributed Execution and Streaming Merge During scan execution, each RegionServer evaluates relational predicates, calculates distances using the pushed down expression, and tracks candidates within a local bounded heap of capacity K, returning only its local top-K rows. On the client, `MergeSortTopNResultIterator` coordinates a streaming K-way merge across all participating region scanners, guaranteeing global ordering without materializing full result sets in memory. Observability is integrated via `ExplainPlanAttributes` and `ExplainTable`. A recognized nearest-neighbor plan replaces the generic `SERVER TOP <n> ROWS SORTED BY [...]` step with `SERVER TOP-<K> BY <distance expression>`, and `MergeSortTopNResultIterator` replaces `CLIENT MERGE SORT` plus a separate `CLIENT LIMIT <K>` with the single step `CLIENT MERGE SORT TOP-<K>`, since the top-K bound is intrinsic to the merge rather than a post-merge truncation. Both steps set the `vectorSearch` flag and the `serverSortAlgo` attribute, and the client step is also appended to the ordered `clientSteps` pipeline, so the disclosure survives into `EXPLAIN (FORMAT JSON)`. At the AST layer, `DistanceFunctionParseNode` exposes metric metadata during semantic analysis. In addition, `PArrayDataType` adds automatic coercion from generic SQL arrays to vector types, allowing applications to pass vector literals or bind parameters as standard arrays without explicit casting. ### Verification and Test Coverage Descriptor extraction, query pattern recognition, constant folding, and error handling (such as missing limits, invalid sort directions, or mismatched dimensions) are tested by `VectorSearchUtilTest` (54 tests). Expression serialization is verified by `DistanceFunctionSerializationTest` (6 tests). Array-to-vector coercion rules are validated across types and dimensions by `PArrayDataTypeVectorCoercionTest` (23 tests). End-to-end cluster execution is proven by `VectorExactSearchIT` (13 tests), which exercises server-side push-down, cross-region streaming merges, all four distance metrics, salted and unsalted primary keys, bind parameters, null handling, query offsets, and explain plan verification on a live mini-cluster. Type round-trip assertions are reinforced in `PDataTypeTest`. --- ## Phase 4: System Catalog and Metadata Extensions (19 files, +1943/−29) ### Catalog Schema and Data Model Approximate vector indexing requires persistent metadata beyond that of standard secondary indexes: the distance metric, indexing algorithm, vector dimensionality, number of IVF posting lists, training sample size, and active centroid generation. This phase establishes the schema model, system catalog storage, and protobuf serialization protocols required to manage vector indexes. In the object model, `PTable` introduces `IndexType.VECTOR_GLOBAL` and exposes accessors for vector metadata: `vectorIndexAlgorithm`, `vectorDistanceMetric`, `vectorDimension`, `vectorIvfLists`, `vectorIvfSampleSize`, and `vectorCentroidGeneration`. `PTableImpl` and its builder handle the serialization, deserialization, and immutability contracts for these properties, while `DelegateTable` provides pass-through delegation. On the wire, `PTable.proto` defines protobuf optional fields 59 through 64 to transport vector index metadata between clients, query services, and RegionServers. ### System Tables and Lifecycle Management Vector metadata is stored across two system tables: 1. `SYSTEM.CATALOG`: Extended with dedicated columns holding vector configuration parameters, managed by `MetaDataEndpointImpl` during table creation (`createTable`), schema lookups (`getTable`), and drops (`dropTable`). Constants for these columns are registered in `PhoenixDatabaseMetaData`. 2. `SYSTEM.VECTOR_CENTROID`: A dedicated system table created to store trained centroid coordinates, generation summaries, and drift scorecards. As defined in `QueryConstants`, the table uses a composite primary key of `(INDEX_NAME, GENERATION_ID, CENTROID_ID)`, ensuring that centroids for a given generation are clustered together in storage. To accommodate these schema additions safely across upgrades, `MetaDataProtocol` advances `MIN_SYSTEM_TABLE_TIMESTAMP_5_4_0` from `MIN_TABLE_TIMESTAMP + 45` to `+ 51`. Both `ConnectionQueryServicesImpl` (for live clusters) and `ConnectionlessQueryServicesImpl` (for unit tests) bootstrap `SYSTEM.VECTOR_CENTROID` alongside existing system tables. Specific error scenarios during vector index creation are cataloged in `SQLExceptionCode`. ### Verification and Test Coverage `VectorSystemCatalogIT` (2 tests) verifies bootstrap creation and column layout of `SYSTEM.VECTOR_CENTROID`. Persistence, retrieval, sentinel row formatting, and multi-generation isolation are tested in `VectorCentroidTableIT` (20 tests). Catalog DDL round-trips and `PTable` accessor accuracy are verified in `VectorIndexIT` (442 lines). Enum serialization and index type predicates are covered by `VectorIndexTypeTest` (13 tests). Catalog metadata serialization is validated in `VectorDataTypeTest`. System table migration and bootstrapping regression suites (`MigrateSystemTablesToSystemNamespaceIT`, `SkipSystemTablesExistenceCheckIT`, and `SystemTablesCreationOnConnectionIT`) are updated and verified against the expanded system table set. --- ## Phase 5: Vector Index DDL and Grammar (12 files, +1514/−28) ### DDL Syntax and Language Grammar Vector indexes are created and modified using explicit SQL DDL syntax: ```sql CREATE VECTOR INDEX idx_name ON table_name (vector_expression) INCLUDE (covered_columns) WITH (metric='L2', algorithm='IVF', lists=1024, sample_size=50000) ``` The grammar in `PhoenixSQL.g` introduces the `create_vector_index_node` production, adds support for `ALTER VECTOR INDEX`, and reserves `VECTOR` as a contextual keyword within DDL statements. The AST captures these specifications in `CreateIndexStatement`, which provides structured getters for the vector expression, metric, algorithm, list count, sample size, and parsed `WITH` properties. ### Compilation and Schema Validation Semantic validation is enforced by `CreateIndexCompiler`. The compiler verifies that: - Exactly one vector expression or column is declared in the index key. - The specified distance metric corresponds to a supported `DistanceMetric` enumeration. - The index algorithm is set to `IVF` (the supported approximate index algorithm). - Posting list counts and training sample sizes fall within valid numeric boundaries. - The target is a physical table; vector indexing over views is explicitly rejected. Once validated, `MetaDataClient` orchestrates index creation. It resolves the base table, confirms the presence and type of referenced columns, constructs the index schema by registering the vector expression as a functional index column, and invokes `MetaDataEndpointImpl` to write the vector metadata and `IndexType.VECTOR_GLOBAL` type marker into `SYSTEM.CATALOG`. ### Verification and Test Coverage DDL compilation rules are thoroughly exercised by `VectorIndexCompilerTest` (18 tests), confirming the acceptance of valid DDL, rejection of invalid metrics or unsupported algorithms, handling of missing parameters, dimension extraction from both physical columns and functional expressions, covered column inclusion, and `IF NOT EXISTS` idempotency. AST construction for complex vector expressions in index keys is validated by `VectorColumnParseTest` (15 tests). End-to-end DDL execution against live clusters is tested in `VectorIndexIT` (432 lines), with additional serialization checks in `VectorIndexTypeTest` and `VectorDataTypeTest`. --- ## Phase 6: IVF Centroid Management (26 files, +8142/−6) ### Clustering Architecture and Metric Adaptation Inverted File (IVF) indexes partition high dimensional space into Voronoi cells using k-means clustering. Centroids define the centers of these cells, serving as partition keys for index rows. This phase implements the training pipeline, storage model, and caching tiers for centroid management across both local and distributed execution environments. Clustering is driven by `KMeansTrainer`, configured via `KMeansConfig`. The algorithm initializes cluster centers using k-means++ seeding, evaluates convergence across configurable iteration limits and tolerance thresholds, and tailors updates to the chosen distance metric: - **Euclidean (L2 / L2_SQUARED):** Standard Lloyd's algorithm computing the arithmetic mean of assigned vectors. - **Cosine:** Spherical k-means, computing the arithmetic mean and subsequently normalizing each centroid vector back to unit length after each iteration. - **Inner Product:** L2-based assignment using arithmetic mean updates. To counteract posting list imbalance caused by data clustering in real world workloads, `KMeansTrainer` performs post training rebalancing. Clusters exceeding variance thresholds are split, and empty clusters are reassigned, with the final effective centroid count and convergence statistics recorded in `KMeansResult`. Cluster health metrics (coefficient of variation, maximum-to-minimum size ratios, empty cluster counts) are computed and serialized by `ClusterSkewMetrics`. ### Distributed Training and Multi-Tier Caching For large tables where in-memory training is infeasible, clustering is distributed across MapReduce via `KMeansDistributedSampler`, `KMeansIterationDriver`, and `KMeansTool`. The distributed sampler runs reservoir sampling per mapper across base table region splits, while the iteration driver orchestrates distributed expectation-maximization iterations across the cluster. Vectors and centroids are serialized through Hadoop execution using `CentroidWritable` and `VectorWritable`, with parameters managed through `PhoenixConfigurationUtil`. Centroid persistence is governed by `CentroidManager`, which interfaces with `SYSTEM.VECTOR_CENTROID`. Centroids are written identified by the current generation, and a sentinal row holds per-generation training statistics. `CentroidManager` also handles lifecycle management, retiring older generations using efficient key-prefix range deletes. To achieve low query latency, `VectorCentroidCache` maintains an LRU cache of centroid arrays on both clients and RegionServers, keyed by `(indexName, generation)`. The cache avoids concurrent cache misses for the same generation. During DDL execution (`CREATE VECTOR INDEX`), `MetaDataClient` samples the base table and executes local training. If the base table is empty or contains fewer non-null vectors than the requested centroid count, `MetaDataClient` records the index metadata in `SYSTEM.CATALOG` while leaving the index in a building state with no centroid generation, deferring training to an asynchronous population build via `ServerBuildIndexCompiler`. Default cache sizes, local training limits, and brute force evaluation thresholds are configured in `QueryServicesOptions`. ### Verification and Test Coverage Algorithmic correctness, convergence guarantees, metric-specific spherical normalization, cluster splitting, and k-means++ reproducibility from fixed random seeds are validated by `KMeansTrainerTest` (15 tests). Centroid CRUD operations, sentinel row encoding, concurrent generation coexistence, and range-delete retirement are tested in `CentroidManagerTest` (18 tests). Cache behavior, including concurrency collapsing, LRU eviction, and generation isolation, is verified by `VectorCentroidCacheTest` (13 tests). Hadoop Writable serialization is covered by `CentroidWritableTest` (3 tests). End-to-end distributed MapReduce training on mini-clusters is tested in `KMeansDistributedTrainerIT` (8 tests). DDL-time training and persistence flows are validated in `VectorCentroidTableIT` (110 lines) and `VectorIndexIT` (99 lines). --- ## Phase 7: Server-Side Write Pipeline for IVF (24 files, +2790/−234) ### Row Key Design and Posting List Layout The performance of IVF indexing depends on organizing index entries into contiguous physical posting lists. The index row key is structured to cluster vectors assigned to the same centroid: ``` [Salt Byte (optional)] [Tenant ID (optional)] [Centroid ID] [Base Table Primary Key] ``` Placing `Centroid ID` immediately after optional salt and tenant prefixes guarantees that all vectors assigned to a given centroid reside in a contiguous lexicographical range. Scanning an IVF partition translates directly into an efficient range scan over that centroid's posting list. ### Coprocessor Write Path and Mutation Lifecycle Index maintenance is coordinated across `IndexMaintainer`, coprocessors, and client paths. During table mutations, the vector expression is evaluated against the row data, its nearest centroid is identified using the active generation retrieved from `VectorCentroidCache`, and the index row key is constructed. On mutable tables, write path maintenance executes inside `IndexRegionObserver.preBatchMutate()`. The coprocessor evaluates both the before-image and after-image of the modified row: - If an insert occurs, it writes a new index row under the assigned centroid. - If an update alters the vector such that it moves across a Voronoi boundary into a different centroid, `IndexMaintainer` generates a paired mutation, a `Delete` for the old centroid's index key and a `Put` for the new centroid's index key. - If covered columns change but the centroid remains identical, an in-place update of the existing index row occurs. For immutable and transactional tables, this assignment is performed on the client. Serialization of vector metadata within index maintainers across RPC boundaries is handled by six new optional fields in `ServerCachingService.proto`. Read repair is integrated through `GlobalIndexChecker`, which validates that an index row's centroid assignment matches the vector value in the base table under the active generation, scheduling repairs for inconsistent entries. Bulk loading and index population are handled by `IndexTool` and `PhoenixIndexImportDirectMapper`, which configure centroid caches on mappers, project vector expressions from base table scans, and emit centroid keyed rows. Supporting updates include `DeleteCompiler`, `PostIndexDDLCompiler`, and `PhoenixIndexBuilder`. In `IndexUtil`, `findVectorColumn()` resolves vector expressions within index schemas, prioritizing functional expression columns before direct column references. In addition, `PVectorFloat` and `PVectorDouble` add sort-order-aware deserialization (`toObject(byte[], SortOrder)`), and `PDataTypeFactory` registers reverse mappings for JDBC metadata. ### Verification and Test Coverage `IndexMaintainerTest` (29 tests) covers vector index maintainer serialization, row key encoding across salted and multi-tenant tables, and centroid reassignment detection. End-to-end write behavior is tested in `VectorIndexWriteIT` (22 tests), verifying inserts, updates, deletes, centroid reassignment, covered column updates, partial row updates, null transitions, and `GlobalIndexChecker` read repairs on a live mini-cluster. Comprehensive write integration is reinforced in `VectorIndexIT` (951 lines). Vector column resolution ordering is tested in `IndexUtilVectorColumnTest` (6 tests). Type conversions and cache interactions under write load are validated in `ColumnInfoToStringEncoderDecoderTest`, `VectorIndexTypeTest`, `VectorDataTypeTest`, and `VectorCentroidCacheTest`. --- ## Phase 8: IVF Query Execution Path (37 files, +7126/−1966) ### Two-Phase Query Execution Architecture IVF similarity queries execute in a coordinated two-phase pipeline between the client and RegionServers, implemented in `VectorIndexScanPlan`: 1. **Client Centroid Probing:** The client retrieves all centroids for the active generation from `VectorCentroidCache` and computes scalar distances against the query vector. It selects the P closest centroids, where P is the probe count (configured via session properties, table defaults, or `VECTOR_PROBE` hints in `HintNode`). `VectorSearchUtil` translates these centroids into multi-range skip-scan filters, expanding them across all salt buckets and aligning them with tenant prefixes. 2. **Server-Side Bounded Scanning:** RegionServers receive the scan request via `NonAggregateRegionScannerFactory`. Instead of scanning the entire index, scanners read only the posting lists corresponding to the selected centroids. Within the RegionServer, `OrderedResultIterator` executes a two-phase scoring loop: - **Coarse Pass:** The scanner evaluates candidate vectors using early terminating bounds (`VectorDistanceUtil.*WithBound`). The bound is dynamically adjusted to the distance of the current K-th worst candidate in the local bounded heap. Vectors whose partial distance exceeds this threshold are discarded immediately, avoiding unnecessary floating point operations. - **Exact Pass:** Candidates that survive the coarse bound are scored completely, updating the local heap. RegionServers return their top candidates (scaled by an oversample factor configured via `ReadOnlyProps`) to the client, where `VectorIndexScanPlan` performs a distributed merge-sort to produce the final top-K result set. ### Optimizer Integration and Test Suite Modernization In `QueryOptimizer`, vector index eligibility is evaluated by comparing distance metrics, vector dimensions, and projected columns against the query's `VectorSearchDescriptor`. When an index matches, `IndexExpressionParseNodeRewriter` maps the base table vector expression to the index's functional column. `ExplainTable` renders execution details, including centroid counts, probe counts, distance metrics, oversample factors, and skip-scan ranges, and reports a vector index as `INDEX <name> VECTOR GLOBAL` alongside the existing `LOCAL`, `GLOBAL`, and `UNCOVERED GLOBAL` kinds. Index candidates that cannot serve the query are rejected with typed reasons drawn from `OptimizerReasons` (`not a nearest-neighbor search`, `distance metric does not match index`, `indexes a different vector column`, `vector dimension does not match index`, and `does not index the query vector expression`), which `EXPLAIN` discloses in `VERBOSE` mode as `/* !INDEX <name> -- <reason> */` lines rather than silen tly dropping the candidate. `ScanUtil` and `IndexTool` are updated to support vector scan attributes. ### Integration Test Refactor This phase also reorganizes the integration test suite, splitting large monolithic test files into specialized classes: `VectorExactSearchIT`, `VectorIndexWriteIT`, `VectorIvfLifecycleIT`, and `VectorIvfQueryIT`, supported by the shared fixture utility `VectorIndexTestUtil`. ### Verification and Test Coverage Query correctness, probe selection, skip-scan filtering, covered versus uncovered projections, multi-region streaming merges, double-precision vectors, oversample variations, and probe hints are validated by `VectorIvfQueryIT` (34 tests), which includes recall assertions against brute force ground truth. Index lifecycle operations (creation, asynchronous builds, drops, rebuilds, and base table mutation interactions) are covered by `VectorIvfLifecycleIT` (14 tests). Exact search pushdown and cross-region merges are verified in `VectorExactSearchIT` (13 tests). Unit tests in `VectorIndexScanPlanTest` (19 tests) validate skip-scan range building, salt expansion, tenant alignment, and centroid ranking. The two phase bounded iterator is tested in `OrderedResultIteratorTest` (7 tests). Extended numerical accuracy and helper tests are added to `DistanceFunctionsTest`, `VectorSearchUtilTest`, `CentroidWritableTest`, and `VectorIndexWriteIT` (22 tests). --- ## Phase 9: Cost-Based Plan Selection and Explain Enhancements (10 files, +1233/−92) ### Cost Model Formulation To decide whether to execute an exact table scan or an approximate IVF index scan, `QueryOptimizer` incorporates a cost model implemented in `CostUtil`. The optimizer estimates execution costs by weighing five factors: 1. **Candidate Cardinality:** Estimated using Phoenix table statistics and guideposts. 2. **Relational Filter Selectivity:** The selectivity of accompanying `WHERE` clause predicates. 3. **Posting List Scan Cost:** Proportional to the probe fraction (P / C), representing the fraction of total centroids C probed by the query. 4. **Base Table Lookup Penalty:** For uncovered queries, the cost of performing random point lookups against the base table for non-indexed columns. 5. **Covered Column Advantage:** The cost savings achieved when all requested columns are included directly within the index. ### Explain Plan Observability `VectorIndexScanPlan` overrides `getExplainPlan`, copying the attributes produced by its superclass into a fresh builder and adding the vector-specific disclosures on top. `ExplainTable` and the JSON renderer then expose: - Which relation was chosen, and why: the scanned table plus the `indexRule` selection label (`cost-based` when the cost model decided it) and, in `VERBOSE` mode, each rejected candidate with its reason. - The probe ratio, as a `CLIENT PROBING <P> OF <C> CENTROIDS (<metric>)` step and as the structured `vectorProbeCount`, `vectorCentroidCount`, and `vectorDistanceMetric` attributes. - Server-side two-phase rescoring, as `SERVER RESCORE TOP-<K> OF <n> CANDIDATES` and the `serverRescoreInfo` attribute, emitted only when the oversample factor exceeds 1.0. - The deferred base table join for an uncovered projection, as a `CLIENT MERGE [<cols>] FOR TOP-<K> ROWS` step spliced in directly after the merge sort, mirrored into `clientSteps` so the JSON pipeline keeps the same ordering, with the columns themselves in `clientMergeColumns`. ### Verification and Test Coverage Plan selection accuracy is verified in `VectorIvfCostIT` (4 tests), confirming that the optimizer selects IVF plans when covered indexes exist, falls back to exact scans when distance metrics diverge or index costs exceed table scans, and prioritizes covered indexes over uncovered alternatives. Cost-driven query execution is reinforced across 301 lines of tests in `VectorIvfQueryIT`, with cost calculation edge cases validated in `VectorSearchUtilTest`. --- ## Phase 10: BSON Vector Extraction (18 files, +1431/−63) ### Document Vector Extraction Semantics Modern architectures often store embeddings within JSON or BSON documents. Phoenix supports this pattern through `BSON_VECTOR_VALUE(document_column, 'dotted.field.path', dimension)`, which extracts a typed `VECTOR(FLOAT, <dim>)` directly from a BSON document. The function expects binary data formatted according to BSON binary subtype 9 using Float32 encoding. Implemented in `BsonVectorValueFunction` and parsed via `BsonVectorValueParseNode` (registered in `ExpressionType`), the function enforces strict data integrity rules: - If the dotted field path does not exist in the document, it returns SQL `NULL`. - If the path matches but has an incorrect BSON binary subtype, invalid floating-point encoding, mismatched byte length, or dimension divergence, it raises an immediate runtime error. ### Functional Indexing over Document Embeddings Extracted vectors integrate directly into Phoenix's secondary indexing engine via functional vector indexes: ```sql CREATE VECTOR INDEX idx ON tbl (BSON_VECTOR_VALUE(doc, 'embedding', 128)) INCLUDE (...) ``` `MetaDataClient` validates that the source column is a valid BSON data type during DDL compilation. On the write path, `IndexMaintainer` (+165 lines) evaluates the extraction against both before-image and after-image document states, correctly handling partial BSON document updates. If the extracted vector shifts across a Voronoi boundary, it emits paired delete/insert mutations to keep posting lists synchronized. During query planning, `QueryOptimizer` and `VectorSearchUtil` match queries referencing `BSON_VECTOR_VALUE` against functional index definitions, while `IndexExpressionParseNodeRewriter` remaps the expression to the index table's functional column. Scan projection and execution are supported by `ProjectionCompiler`, `NonAggregateRegionScannerFactory`, and `BaseScannerRegionObserverConstants.BSON_VECTOR_VALUE_FUNCTION`, with wire metadata carried in `ServerCachingService.proto`. `ProjectionCompiler` registers the pushed down extractions on the statement context, so they join `BSON_VALUE` in the shared `SERVER BSON PROJECTION <n>` clause with one indented detail line pe r server evaluated expression. ### Verification and Test Coverage Function evaluation, missing path handling, subtype validation, dimension mismatch enforcement, and expression serialization are tested by `BsonVectorValueFunctionTest` (16 tests). Parsing of BSON vector expressions in DDL statements is verified in `VectorColumnParseTest`. End-to-end integration is proven in `VectorIndexIT` (474 lines), validating index creation, bulk population, query execution, covered column retrieval, and document update-driven index maintenance. --- ## Phase 11: Single-Cell Storage Support for Vector Indexes (7 files, +731/−75) ### Storage Layout Compatibility Phoenix supports two storage layouts for immutable tables: standard `ONE_CELL_PER_COLUMN` and `SINGLE_CELL_ARRAY_WITH_OFFSETS`. In the single cell layout, all column values within a column family are packed into a single HBase cell alongside an offset array. This format optimizes disk footprints and read performance for dense wide tables. Vector indexes must function identically across both storage models. ### Maintainer Adaptations and Transcoding In `IndexMaintainer`, column expression evaluation was refactored to recognize and unwrap `SingleCellColumnExpression` wrappers. This enables the write pipeline to extract vector values from packed single cell cells during mutation processing. During partial updates of covered columns, the maintainer preserves vector values across row reconstruction phases. Furthermore, sort-order transcoding is normalized when reading vector columns from single cell tables, ensuring consistent binary keys regardless of physical storage format. Metadata propagation in `MetaDataClient` and `MetaDataEndpointImpl` was updated to ensure that immutable storage scheme attributes are preserved during vector index creation and administrative operations. ### Verification and Test Coverage Verification in `VectorIndexWriteIT` (271 lines) confirms that insert, update, delete, and centroid reassignment mutations generate identical index row keys under both `ONE_CELL_PER_COLUMN` and `SINGLE_CELL_ARRAY_WITH_OFFSETS`. Query consistency across storage schemes is proven in `VectorIvfQueryIT` (139 lines), verifying that IVF search recall is unaffected by the underlying table format. Unit tests in `IndexMaintainerTest` (215 lines) validate expression unwrapping and row key generation, while DDL persistence is verified in `VectorIndexIT` (70 lines). --- ## Phase 12: Adaptive Probing and Hybrid Filtering (15 files, +1675/−44) ### Adaptive Probe Expansion When similarity queries include selective relational predicates, scanning only the initial P centroids can result in candidate starvation, a state where relational filters discard most candidates, returning fewer than the requested K results. Adaptive probing resolves this within `VectorIndexScanPlan`. The query begins by scanning posting lists for the initial P closest centroids. If post-filtering leaves fewer than K candidates, the plan initiates dynamic probe expansion: 1. It computes distances to the remaining centroids to identify the next closest batch. 2. It constructs new multi-range skip-scan filters for these posting lists. 3. It dispatches supplementary scans to the RegionServers. 4. It merges newly discovered candidates into the existing candidate pool using `MergeSortTopNResultIterator`. This cycle repeats until K candidates are accumulated, all viable centroids are scanned, or a configured maximum probe limit is reached. The maximum probe threshold can be specified per-query using the `VECTOR_MAX_PROBE` hint in `HintNode`, or globally via `QueryServices.VECTOR_MAX_PROBE_LIMIT` and `QueryServicesOptions`. ### Hybrid Filter-First Plan Selection In scenarios where a query includes a highly selective scalar predicate backed by an existing standard secondary index, scanning an IVF index even with adaptive probing can be suboptimal. For these workloads, `QueryOptimizer` introduces `FilterFirstIndexPlan` (supported by `ScanPlan` and `DelegateQueryPlan`). The filter-first plan executes in reverse order: 1. It uses the scalar secondary index to locate rows that satisfy the selective predicate. 2. It retrieves the vector columns for only those matching rows. 3. It computes exact vector distances and ranks the candidates. The optimizer compares the estimated cost of `FilterFirstIndexPlan` against `VectorIndexScanPlan`, leveraging `VectorSearchUtil` for filter selectivity estimation and `BsonValueFunction` for document predicate analysis. ### Verification and Test Coverage Adaptive probe expansion, candidate starvation recovery, maximum probe bounds, filter-first plan selection, and covered/uncovered column interactions are verified in `VectorIndexIT` (736 lines). Unit tests in `VectorIndexScanPlanTest` (95 lines) validate batch range generation, cumulative range merging, and limit enforcement. Predicate extraction for filter-first plans is verified in `BsonValueFunctionTest` (95 lines), and selectivity estimation routines are tested in `VectorSearchUtilTest` (119 lines). --- ## Phase 13: Drift Scorecarding and Background Rebuild (35 files, +6138/−144) ### Continuous Health Monitoring and Scorecard Storage Over time, shifts in vector data distributions cause cluster imbalance. Some posting lists grow disproportionately large while others become sparse, degrading scan efficiency and reducing search recall. Phoenix addresses this with an inline drift scorecarding and automated rebuild framework. Index health is tracked in `SYSTEM.VECTOR_CENTROID` as durable scorecard rows stored alongside the centroids they monitor (one row per centroid per generation, plus a sentinel summary row). The write path updates the scorecard inline. Because `IndexMaintainer` already inspects before-images and after-images to derive centroid assignments, it calculates posting list population deltas with negligible overhead. To prevent write amplification on the system catalog, RegionServers buffer deltas in memory using `ScorecardAccumulator`. Inside `IndexRegionObserver`, accumulators aggregate mutation counts across configurable time windows (`QueryServices.VECTOR_SCORECARD_FLUSH_INTERVAL_MS`), periodically flushing deltas to `SYSTEM.VECTOR_CENTROID` using atomic aggregating upserts (`ScorecardRow`). ### Counter Reconciliation and Drift Assessment Because node crashes or aborted transactions can introduce minor discrepancies in in-memory counters, Phoenix provides scheduled reconciliation. `VectorScorecardReconcileTask` (`TaskType.VECTOR_SCORECARD_RECONCILE` in `PTable`), registered in `TaskRegionObserver`, executes a lightweight grouped count scan along the index's leading key column. This scan recalculates exact posting list counts, correcting counter drift. Reconciliation runs on a background schedule (`QueryServices.VECTOR_SCORECARD_RECONCILE_INTERVAL_MS`) and is executed before any rebuild decision. Index drift is evaluated by `VectorIndexScorecard` against four administrative thresholds: 1. **Posting List Skew Ratio:** Ratio of the largest posting list to the median list size. 2. **Coefficient of Variation (CV):** Standard deviation of cluster sizes divided by the mean. 3. **Empty Centroid Fraction:** Percentage of centroids with zero assigned vectors. 4. **Reassignment Rate:** Rate at which document updates shift vectors across centroid boundaries. Assessment verdicts and diagnostic summaries are recorded in `GenerationSummary` within the sentinel row of `SYSTEM.VECTOR_CENTROID`. ### Zero-Downtime Rebuild and Migration When drift thresholds are crossed, or upon manual invocation via `MetaDataClient`, `VectorIndexRebuildTask` (`TaskType.VECTOR_INDEX_REBUILD`) coordinates an asynchronous background rebuild: 1. **Model Training:** It samples the base table and trains a new centroid generation (G_new). 2. **Data Migration:** Using `CentroidManager`, it scans the index and rewrites entries under the new centroid model using atomic paired delete/insert mutations. During migration, queries continue executing against the active generation (G_old). 3. **Generation Promotion:** Once migration is complete, the task updates the index's active generation to G_new in `SYSTEM.CATALOG`, initializes the new generation's scorecard, and retires G_old. 4. **Garbage Collection:** Obsolete centroid rows and scorecard data for G_old are removed via key-prefix range deletes. During rebuild operations, query probe counts can be dynamically adjusted according to configured rebuild probe policies (`QueryServices.VECTOR_REBUILD_PROBE_POLICY` and factor). `EXPLAIN` discloses that adjustment so a plan captured mid-rebuild is not mistaken for steady state: the `EXPAND` policy annotates the probe step as `CLIENT PROBING <P> OF <C> CENTROIDS (<metric>) (REBUILD IN PROGRESS: EXPANDED FROM <P_base>)`, and the `EXACT` policy replaces it with `CLIENT EXACT VECTOR EVALUATION OVER <C> CENTROIDS (REBUILD IN PROGRESS)`. Constants and DDL schemas for scorecard columns are registered in `PhoenixDatabaseMetaData` and `QueryConstants`. ### Verification and Test Coverage Scorecard evaluation logic, CV calculations, empty-cluster detection, and verdict generation are verified in `VectorIndexScorecardTest` (11 tests). Reconciliation accuracy against deliberately corrupted counters and task scheduling are tested in `VectorScorecardReconcileTaskTest` (5 tests). Centroid migration, generation promotion, and range retirement are validated in `CentroidManagerTest` (298 lines). Migration correctness and generation coexistence during live rebuilds are proven in `VectorIvfLifecycleIT` (362 lines). Scorecard persistence and reconciliation integration are verified in `VectorCentroidTableIT` (589 lines). End-to-end drift detection, delta accumulation, flush verification, and automated background rebuilds on live clusters are tested in `VectorIndexIT` (1384 lines). Additional probe policy and scorecard assertions are covered in `VectorIndexScanPlanTest`, `VectorColumnIT`, and `VectorIndexWriteIT`. --- ## Phase 14: Compatibility and Version Validation (7 files, +252/−22) ### Cross-Version Compatibility and Protocol Gating In distributed environments running rolling upgrades or mixed client deployments, older clients that lack vector index awareness must be prevented from executing invalid plans or corrupting index metadata. Conversely, newer clients must handle connections to older server clusters cleanly. This phase implements strict version gating: - **Client-Side Validation:** `ConnectionQueryServices` and `ConnectionQueryServicesImpl` inspect cluster capabilities during connection initialization. Clients reject vector index metadata if the server reports a version below the vector support threshold. - **DDL Gating:** In `MetaDataClient`, attempts to create a vector index against a cluster below the minimum supported version fail immediately with an explicit compatibility error. - **Server-Side Protection:** `IndexRegionObserver` inspects index maintainers during mutation processing. If a maintainer contains vector specifications on a RegionServer version that predates vector support, it logs a warning and skips vector-specific handling rather than failing unexpectedly. ### Wire Protocol Backward Compatibility Protocol buffer serialization in `IndexMaintainer` (+51 lines) adheres to strict backward-compatibility rules. When serializing metadata for older peers, vector-specific fields are omitted to prevent deserialization errors. During deserialization, missing optional vector fields are interpreted as standard non-vector secondary indexes, ensuring unupgraded nodes continue normal operation without interruption. ### Verification and Test Coverage Version validation behavior is tested in `VectorIndexIT` (89 lines), verifying that older client simulations reject vector metadata, older server simulations fail vector DDL with descriptive errors, and protobuf serialization round-trips maintain wire compatibility. Protocol compatibility and version-gated serialization are validated in `VectorIndexTypeTest` (67 lines). --- ## Test Verification and Validation All test suites were executed against the branch head (`582ecfaece`) on macOS (Darwin 25.5.0, Apple Silicon) using JDK 11 (Temurin 11.0.30) for baseline verification and JDK 21 (Temurin 21.0.12.1) for Panama SIMD validation. The test counts and pass/fail outcomes below are from a re-execution after the branch was rebased onto base `f590b9d9e0`. The integration timings are from that run too, measured serially with one `mvn verify` per class. The unit timings are indicative only: the unit suite runs 8 reused forks in parallel, so a class that lands in an already warm JVM reports a fraction of the time the same class takes in a cold one. The JDK 21 Panama overlay figures in the final subsection carry over from the pre-rebase verification, as the rebase touched no distance kernel. Because the vector attributes are additive to `ExplainPlanAttributes`, they are also exercised by the `EXPLAIN` regression suites that ship with the base commit. `ExplainPlanTest` compares a complete normalized JSON attribute tree for each of its 111 corpus cases, so every vector attribute is asserted, in its unset form, across all of them; `QueryOptimizerTest`, `QueryCompilerTest`, `ExplainOptionsParserTest`, and `ExplainJsonOutputTest` cover the index selection rule and rejected-candidate plumbing that vector index rejections feed into. Those five suites total 397 tests and pass, with 3 pre-existing skips. ### Unit Test Execution The unit test suite completed with **419 passed tests, 0 failures, and 0 errors**, executed with the project default of 8 parallel forks (`numForkedUT`). One unrelated class, `PhoenixStatsCacheLoaderTest`, can stall under that level of fork contention: `testStatsBeingAutomaticallyRefreshed` waits on an untimed `CountDownLatch` for a timing-dependent stats cache refresh. It is pre-existing, touches no vector or `EXPLAIN` code, and passes in under a second when its class is run alone. | Test Class | Tests | Execution Time | Primary Test Coverage | |---|---|---|---| | `VectorSearchUtilTest` | 54 | 5.4s | Plan descriptor extraction, metric resolution, constant folding, edge cases | | `VectorDataTypeTest` | 43 | 0.2s | Codec round-trips, sort orders, nullability, dimension validation | | `PDataTypeTest` | 38 | 0.5s | Type registration, global type-iteration stability | | `IndexMaintainerTest` | 29 | 3.4s | Row key encoding, centroid reassignments, single-cell unwrapping | | `DistanceFunctionsTest` | 27 | 2.6s | End-to-end evaluation, dimensional compatibility, accuracy against baselines | | `PArrayDataTypeVectorCoercionTest` | 23 | 0.01s | Array-to-vector coercion rules across element types and dimensions | | `DistanceOperatorParseTest` | 20 | 1.5s | Infix operator parsing, precedence, AST normalization | | `VectorColumnParseTest` | 19 | 1.9s | Column definitions, DDL grammar, constraint parsing | | `VectorIndexScanPlanTest` | 19 | 2.9s | Skip-scan range building, salt expansion, tenant alignment, ranking | | `CentroidManagerTest` | 18 | 1.5s | Generation CRUD, sentinel rows, lifecycle retirement by range delete | | `VectorIndexCompilerTest` | 18 | 3.2s | DDL compilation, metric validation, algorithm validation, covered columns | | `BsonVectorValueFunctionTest` | 16 | 0.2s | BSON path extraction, binary subtype 9 validation, dimension checks | | `KMeansTrainerTest` | 15 | 0.7s | Convergence detection, cluster splitting, spherical k-means, seeding | | `VectorCentroidCacheTest` | 13 | 2.2s | Hit/miss ratios, LRU eviction, concurrency collapsing, generation isolation | | `VectorIndexTypeTest` | 13 | 0.1s | IndexType serialization, ordinal stability, version-gated serde | | `VectorIndexScorecardTest` | 11 | 2.2s | Scorecard evaluation, skew ratios, CV computation, reassignment rates | | `OrderedResultIteratorTest` | 7 | 2.5s | Two-phase scoring, early-termination bounding, heap limits | | `DistanceFunctionSerializationTest` | 6 | 0.03s | Serde round-trips through Phoenix expression serialization framework | | `IndexUtilVectorColumnTest` | 6 | 1.2s | Vector column resolution ordering (functional expressions vs. direct refs) | | `CoerceExpressionTest` | 5 | 0.06s | Vector-to-vector and array coercion expression rules | | `VectorScorecardReconcileTaskTest` | 5 | 0.9s | Counter reconciliation accuracy under corruption, task scheduling | | `VectorDistanceUtilTest` | 3 | 0.01s | Low-level scalar kernels (L2, cosine, dot product) on packed buffers | | `BsonValueFunctionTest` | 3 | 2.5s | BSON predicate extraction for filter-first plan selection | | `CentroidWritableTest` | 3 | 0.01s | Hadoop Writable serialization round-trips for centroid data | | `ColumnInfoToStringEncoderDecoderTest` | 3 | 0.1s | ColumnInfo serialization with vector data types | | `IndexUtilTest` | 2 | 1.4s | Secondary index utility constants and column lookup validation | ### Integration Test Execution The integration test suite completed with **214 passed tests, 0 failures, 0 errors, and 1 skipped**, run strictly serially with one `mvn verify` invocation per class (total wall time approximately 43 minutes): | Test Class | Tests | Execution Time | Primary Test Coverage | |---|---|---|---| | `VectorIndexIT` | 63 | 6m 42s | End-to-end DDL, writes, queries, scorecarding, and background rebuilds | | `VectorIvfQueryIT` | 34 | 53s | IVF execution, multi-region merge, covered projections, recall verification | | `VectorIndexWriteIT` | 22 | 1m 39s | Write pipeline, centroid reassignment, single-cell storage, read repairs | | `VectorCentroidTableIT` | 20 | 5s | System table CRUD, generation isolation, sentinel rows, scorecard persistence | | `SystemTablesCreationOnConnectionIT` | 15 (1 skipped) | 11m 20s | System table bootstrap validation (includes `SYSTEM.VECTOR_CENTROID`) | | `VectorIvfLifecycleIT` | 14 | 1m 34s | Index build, async population, drop, rebuild, generation coexistence | | `VectorExactSearchIT` | 13 | 0.2s | Server push-down, cross-region streaming merge, offsets, bind variables | | `VectorColumnIT` | 12 | 25s | Vector column DDL, upsert binding, JDBC metadata, constraint enforcement | | `KMeansDistributedTrainerIT` | 8 | 1m 9s | Distributed MapReduce k-means training on mini-cluster | | `VectorIvfCostIT` | 4 | 39s | Cost-based plan selection (IVF vs. exact scan vs. covered index preference) | | `MigrateSystemTablesToSystemNamespaceIT` | 4 | 5m 57s | Namespace migration containing new system tables | | `SkipSystemTablesExistenceCheckIT` | 3 | 1m 10s | Fast connection bootstrap without redundant system table checks | | `VectorSystemCatalogIT` | 2 | 0.2s | Bootstrap creation and column verification of `SYSTEM.VECTOR_CENTROID` | The single skipped test corresponds to a pre-existing `@Ignore` annotation in `SystemTablesCreationOnConnectionIT`, which is unrelated to vector indexing. ### JDK 21 SIMD Overlay Verification The `java21` Maven profile introduces a multi-release JAR overlay with Panama Vector API distance kernels. Because Phoenix's default compile target is JDK 11, this overlay is maintained in `src/main/java21` and validated through a one-off verification procedure: 1. **Compilation:** The overlay compiles cleanly under JDK 21 with `--release 21 --add-modules jdk.incubator.vector`. 2. **API Parity:** The Java 21 `VectorDistanceUtil` implementation matches the public method signatures of the Java 11 base class, adhering to the multi-release JAR specification. 3. **Algorithmic Equivalence:** 104 tests across five suites (`DistanceFunctionsTest`, `VectorDataTypeTest`, `BsonVectorValueFunctionTest`, `KMeansTrainerTest`, and `VectorDistanceUtilTest`) were re-run under JDK 21 with the Panama overlay active on the classpath. All tests passed, confirming that the SIMD-accelerated kernels produce results identical to the scalar reference paths. ### Testing Notes The test harness incorporates several architectural testing standards: - **Deterministic Geometric Fixtures:** Rather than relying on non-deterministic random data, integration tests utilize geometric fixtures with mathematically predetermined centroid placements and known cluster memberships (`VectorIndexTestUtil`). This allows exact recall and Voronoi boundary transitions to be verified deterministically. - **Independent Ground-Truth Comparison:** Exact nearest-neighbor searches are validated against independently computed brute-force rankings to prevent circular verification. - **Exhaustive Code Coverage:** Across 39 modified and new test files containing 622 `@Test` annotations, all new execution paths (including failure modes, boundary transitions, and error codes) maintain full coverage without placeholder assertions or ignored test methods. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
