[ 
https://issues.apache.org/jira/browse/PHOENIX-8009?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18119717#comment-18119717
 ] 

Andrew Kyle Purtell commented on PHOENIX-8009:
----------------------------------------------

For vector graph indexes I have two proposals that have different pros and cons.

HNSW Graph-Based Vector Indexes (JVector + MOB Storage): 
https://gist.github.com/apurtell/e4668cddf949c43955d4a919489d4228 

HNSW Graph-Based Vector Indexes (Cell-Based Implementation): 
https://gist.github.com/apurtell/0baabc65a7234b8171455905acb34bd9

> Vector Indexes Phase 2
> ----------------------
>
>                 Key: PHOENIX-8009
>                 URL: https://issues.apache.org/jira/browse/PHOENIX-8009
>             Project: Phoenix
>          Issue Type: Sub-task
>            Reporter: Andrew Kyle Purtell
>            Assignee: Andrew Kyle Purtell
>            Priority: Major
>
> The IVF algorithm delivered in PHOENIX-7998 provides effective approximate 
> nearest-neighbor search by partitioning high dimensional space into Voronoi 
> cells, but its cluster-centric design imposes recall limitations under high 
> dimensional skew and requires periodic centroid retraining as data 
> distributions shift. Hierarchical Navigable Small World (HNSW) graph indexes 
> can deliver higher recall at equivalent query latency and achieve two-fold to 
> five-fold recall improvements over IVF at equivalent latency budgets by 
> introducing a graph-based ANN algorithm alongside the existing IVF algorithm. 
> This capability targets similarity search queries where high recall is 
> paramount, update heavy workloads where IVF centroid drift degrades quality 
> between rebuild cycles, and moderate scale vector collections where the 
> memory footprint remains tractable.
> Algorithmic foundations draw from the HNSW construction algorithm formalized 
> by Malkov and Yashunin in 2018 combined with the Vamana robust pruning 
> heuristic from the DiskANN work by Subramanya et al. in 2019, adapted to 
> Phoenix's distributed coprocessor execution model and HBase storage 
> semantics. Integrated vector quantization draws from scalar quantization and 
> product quantization formalized by Jégou, Douze, and Schmid in 2011, fused 
> directly into the graph storage layout and search pipeline to minimize block 
> cache pressure and I/O amplification during graph traversal.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to