[ 
https://issues.apache.org/jira/browse/AIRAVATA-3970?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18066237#comment-18066237
 ] 

Jayanth Vennamreddy commented on AIRAVATA-3970:
-----------------------------------------------

[~smarru] I've updated the description to generalize schema design for multiple 
MD databases while also keeping implementation focused and scoped for the GSoC 
timeline

> [GSoC] Adaptive Metadata Indexing and ATLAS Integration in Airavata Data 
> Catalog
> --------------------------------------------------------------------------------
>
>                 Key: AIRAVATA-3970
>                 URL: https://issues.apache.org/jira/browse/AIRAVATA-3970
>             Project: Airavata
>          Issue Type: Improvement
>            Reporter: Jayanth Vennamreddy
>            Priority: Major
>              Labels: gsoc, gsoc2026, mentor
>
> Background:
> Apache Airavata’s data catalog provides structured metadata storage but 
> currently lacks workload-aware optimization strategies and efficient support 
> for filter-heavy scientific metadata queries.
> With potential integration of external scientific datasets such as ATLAS (a 
> molecular dynamics database containing ~1,900 protein simulations with rich 
> structural, domain, and MD metrics metadata), the catalog must support more 
> advanced retrieval patterns, including:
>  - Key-based lookups by PDB chain
>  - Multi-field metadata filtering (organism, domain classification, 
> resolution, MD metrics)
>  - Batch metadata retrieval following similarity searches
>  - Scalable performance under increasing dataset size (10k–100k+ records)
> Currently, metadata retrieval mechanisms are optimized for structured storage 
> but do not incorporate workload-aware indexing or filter-efficient retrieval 
> for domain-rich scientific data. While ATLAS serves as an initial integration 
> target, the schema and indexing framework will be designed to support 
> heterogeneous molecular dynamics databases such as mdCATH, GPCRmd, and 
> MemProtMD, which differ in classification systems, primary identifiers, and 
> metadata structures.
> Problem:
> Static indexing strategies do not scale effectively for scientific metadata 
> workloads where query patterns vary across research users. Additionally, 
> current APIs are optimized primarily for single-key access and do not support 
> efficient bulk or filter-driven retrieval.
> As Airavata evolves to support protein-scale metadata and 
> similarity-search-driven workflows, improvements in schema design, indexing 
> strategy, and retrieval efficiency become necessary.
> Proposed Work:
> 1. Design and implement a normalized, extensible metadata schema in 
> Airavata’s data catalog capable of representing protein simulation metadata 
> from multiple MD databases (e.g., ATLAS, mdCATH, GPCRmd, MemProtMD), with 
> support for multi-value classification fields.
> 2. Implement a metadata ingestion pipeline to import ATLAS protein records 
> (~1,900 entries) into the catalog and validate correctness.
> 3. Add query telemetry instrumentation to metadata APIs to capture:
>    - Filter predicates used
>    - Query latency
>    - Result set size
>    - Field access frequency
> 4. Based on observed workload patterns, implement an index optimization 
> module that:
>    - Identifies high-frequency filter fields
>    - Automatically creates appropriate secondary or composite indexes based 
> on observed workload thresholds, with controlled evaluation before activation.
>    - Benchmarks performance before and after index creation
> 5. Implement a batch metadata retrieval API optimized for 
> similarity-search-driven workflows, enabling efficient bulk fetch of protein 
> metadata records.
> 6. Evaluate performance under synthetic scaling (10k–100k records) to measure 
> query latency improvements and indexing overhead.
> Expected Outcomes:
>  - ATLAS metadata fully represented in Airavata’s data catalog with proper 
> schema support.
>  - Schema validation against at least one additional MD database to ensure 
> generality beyond ATLAS.
>  - Telemetry-driven index optimization workflow implemented.
>  - Demonstrated reduction in metadata query latency for filter-heavy queries.
>  - Support for bulk metadata retrieval following similarity searches.
>  - Benchmarks and documentation demonstrating scalability improvements.
> This issue will serve as the tracking issue for the GSoC 2026 proposal and 
> scope discussion.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to