[
https://issues.apache.org/jira/browse/PHOENIX-8009?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18119717#comment-18119717
]
Andrew Kyle Purtell commented on PHOENIX-8009:
----------------------------------------------
For vector graph indexes I have two proposals that have different pros and cons.
HNSW Graph-Based Vector Indexes (JVector + MOB Storage):
https://gist.github.com/apurtell/e4668cddf949c43955d4a919489d4228
HNSW Graph-Based Vector Indexes (Cell-Based Implementation):
https://gist.github.com/apurtell/0baabc65a7234b8171455905acb34bd9
> Vector Indexes Phase 2
> ----------------------
>
> Key: PHOENIX-8009
> URL: https://issues.apache.org/jira/browse/PHOENIX-8009
> Project: Phoenix
> Issue Type: Sub-task
> Reporter: Andrew Kyle Purtell
> Assignee: Andrew Kyle Purtell
> Priority: Major
>
> The IVF algorithm delivered in PHOENIX-7998 provides effective approximate
> nearest-neighbor search by partitioning high dimensional space into Voronoi
> cells, but its cluster-centric design imposes recall limitations under high
> dimensional skew and requires periodic centroid retraining as data
> distributions shift. Hierarchical Navigable Small World (HNSW) graph indexes
> can deliver higher recall at equivalent query latency and achieve two-fold to
> five-fold recall improvements over IVF at equivalent latency budgets by
> introducing a graph-based ANN algorithm alongside the existing IVF algorithm.
> This capability targets similarity search queries where high recall is
> paramount, update heavy workloads where IVF centroid drift degrades quality
> between rebuild cycles, and moderate scale vector collections where the
> memory footprint remains tractable.
> Algorithmic foundations draw from the HNSW construction algorithm formalized
> by Malkov and Yashunin in 2018 combined with the Vamana robust pruning
> heuristic from the DiskANN work by Subramanya et al. in 2019, adapted to
> Phoenix's distributed coprocessor execution model and HBase storage
> semantics. Integrated vector quantization draws from scalar quantization and
> product quantization formalized by Jégou, Douze, and Schmid in 2011, fused
> directly into the graph storage layout and search pipeline to minimize block
> cache pressure and I/O amplification during graph traversal.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)