[
https://issues.apache.org/jira/browse/AIRAVATA-3970?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18066237#comment-18066237
]
Jayanth Vennamreddy commented on AIRAVATA-3970:
-----------------------------------------------
[~smarru] I've updated the description to generalize schema design for multiple
MD databases while also keeping implementation focused and scoped for the GSoC
timeline
> [GSoC] Adaptive Metadata Indexing and ATLAS Integration in Airavata Data
> Catalog
> --------------------------------------------------------------------------------
>
> Key: AIRAVATA-3970
> URL: https://issues.apache.org/jira/browse/AIRAVATA-3970
> Project: Airavata
> Issue Type: Improvement
> Reporter: Jayanth Vennamreddy
> Priority: Major
> Labels: gsoc, gsoc2026, mentor
>
> Background:
> Apache Airavata’s data catalog provides structured metadata storage but
> currently lacks workload-aware optimization strategies and efficient support
> for filter-heavy scientific metadata queries.
> With potential integration of external scientific datasets such as ATLAS (a
> molecular dynamics database containing ~1,900 protein simulations with rich
> structural, domain, and MD metrics metadata), the catalog must support more
> advanced retrieval patterns, including:
> - Key-based lookups by PDB chain
> - Multi-field metadata filtering (organism, domain classification,
> resolution, MD metrics)
> - Batch metadata retrieval following similarity searches
> - Scalable performance under increasing dataset size (10k–100k+ records)
> Currently, metadata retrieval mechanisms are optimized for structured storage
> but do not incorporate workload-aware indexing or filter-efficient retrieval
> for domain-rich scientific data. While ATLAS serves as an initial integration
> target, the schema and indexing framework will be designed to support
> heterogeneous molecular dynamics databases such as mdCATH, GPCRmd, and
> MemProtMD, which differ in classification systems, primary identifiers, and
> metadata structures.
> Problem:
> Static indexing strategies do not scale effectively for scientific metadata
> workloads where query patterns vary across research users. Additionally,
> current APIs are optimized primarily for single-key access and do not support
> efficient bulk or filter-driven retrieval.
> As Airavata evolves to support protein-scale metadata and
> similarity-search-driven workflows, improvements in schema design, indexing
> strategy, and retrieval efficiency become necessary.
> Proposed Work:
> 1. Design and implement a normalized, extensible metadata schema in
> Airavata’s data catalog capable of representing protein simulation metadata
> from multiple MD databases (e.g., ATLAS, mdCATH, GPCRmd, MemProtMD), with
> support for multi-value classification fields.
> 2. Implement a metadata ingestion pipeline to import ATLAS protein records
> (~1,900 entries) into the catalog and validate correctness.
> 3. Add query telemetry instrumentation to metadata APIs to capture:
> - Filter predicates used
> - Query latency
> - Result set size
> - Field access frequency
> 4. Based on observed workload patterns, implement an index optimization
> module that:
> - Identifies high-frequency filter fields
> - Automatically creates appropriate secondary or composite indexes based
> on observed workload thresholds, with controlled evaluation before activation.
> - Benchmarks performance before and after index creation
> 5. Implement a batch metadata retrieval API optimized for
> similarity-search-driven workflows, enabling efficient bulk fetch of protein
> metadata records.
> 6. Evaluate performance under synthetic scaling (10k–100k records) to measure
> query latency improvements and indexing overhead.
> Expected Outcomes:
> - ATLAS metadata fully represented in Airavata’s data catalog with proper
> schema support.
> - Schema validation against at least one additional MD database to ensure
> generality beyond ATLAS.
> - Telemetry-driven index optimization workflow implemented.
> - Demonstrated reduction in metadata query latency for filter-heavy queries.
> - Support for bulk metadata retrieval following similarity searches.
> - Benchmarks and documentation demonstrating scalability improvements.
> This issue will serve as the tracking issue for the GSoC 2026 proposal and
> scope discussion.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)