alexanderbianchi opened a new issue, #24:
URL: https://github.com/apache/datafusion-iceberg/issues/24

   ## Motivation
   
   The current catalog integration combines long-lived catalog discovery, 
cached  table providers, and fresh metadata loading during physical planning. 
These lifetimes do not align well with query-scoped authentication or 
consistent schema resolution.
   
   @DerGut's https://github.com/apache/iceberg-rust/pull/3000 addresses some of 
these issues in connecting the `SessionCatalog` work in `iceberg-rust` 
https://github.com/apache/iceberg-rust/pull/2920. Discussions there are 
relevant for this issue.
   
   This issue is intended to discuss the current limitations and general 
direction for the catalog integration.  I’ll create an EPIC with separate 
follow-up issues for each component.
   
   ### 1. Catalog contents and providers are cached for the provider’s lifetime
   
    The catalog provider eagerly stores every schema:
   
    ```rust
      pub struct IcebergCatalogProvider {
          schemas: HashMap<String, Arc<dyn SchemaProvider>>,
      }
    ```
    Each schema provider eagerly stores every table provider:
    ```rust
      pub(crate) struct IcebergSchemaProvider {
          catalog: Arc<dyn Catalog>,
          namespace: NamespaceIdent,
          tables: Arc<DashMap<String, Arc<IcebergTableProvider>>>,
      }
    ```
   Construction lists namespaces, then lists tables in each discovered 
namespace, effectively enumerating the whole catalog.
    ```rust
      client.list_namespaces(None).await?
        ...
      client.list_tables(&namespace).await?
    ```
   
    The same Arc<IcebergTableProvider> is returned to every query:
   
    ```rust
      Ok(self
          .tables
          .get(name)
          .map(|entry| entry.value().clone() as Arc<dyn TableProvider>))
    ```
   
    Consequences:
   
    - tables created or dropped outside this provider are not reflected
    - startup requires catalog-wide listing permissions
    - initialization cost scales with the entire catalog rather than the query
    - providers cannot naturally be scoped to a tenant or authenticated request
    - querying one table requires initializing providers for unrelated tables.
   
   ### 2. Table resolution and scan planning happen too late
   
   The table provider caches its schema during construction, but reloads the 
table during scan and insert planning. If the table changes between those 
steps, planning can combine the cached schema with newer metadata. This also 
repeats catalog requests for a table already loaded during resolution.
   File discovery and scan planning also happen during execution, making it 
harder to expose partitioned work to DataFusion or distribute that work without 
repeating planning.
   
   ### 3. Remote catalog operations do not fit the synchronous provider 
lifecycle
   
   Remote Iceberg operations are asynchronous, while parts of DataFusion’s 
catalog interface are synchronous. The current integration handles discovery by 
eagerly loading and caching catalog contents during construction. For 
mutations, SchemaProvider::register_table() and deregister_table() instead 
block on remote create/drop operations. This occupies a blocking-pool thread 
while awaiting remote I/O and makes cancellation and runtime behavior harder to 
reason about
   
   These are two consequences of the same mismatch: remote catalog access needs 
an explicit asynchronous boundary, with resolved providers serving subsequent 
planning lookups.
   
   ## What a proper integration should provide
   
   ### Asynchronous, query-scoped resolution
   
   - Load only referenced tables, resolving repeated references once.
   - Avoid requiring catalog-wide listing permissions to query a known table.
   - Reuse the catalog client across queries and resolved providers throughout 
each query.
   - Keep catalog resolution, scan planning, and execution as separate stages.
   
   #### AsyncCatalogProvider helps, but context must be bound explicitly
   
   [[DataFusion’s asynchronous catalog 
traits](https://docs.rs/datafusion/latest/datafusion/catalog/trait.AsyncCatalogProvider.html)](https://docs.rs/datafusion/latest/datafusion/catalog/trait.AsyncCatalogProvider.html)
 support resolving references into query-local providers before planning. 
However, `resolve()` receives only `SessionConfig`, and the lower-level 
schema/table lookups receive no session context. The integration therefore 
needs an explicit way to bind trusted request context before resolution.
   
   ### Request context bound before catalog access
   
   - Let the embedding application supply trusted identity and credentials 
before resolution.
   - Preserve that context through catalog operations, including commit refresh 
and retries.
   - Isolate concurrent queries, define a stable session identity, and fail 
early when required context is missing.
   - Keep catalog credentials out of SQL-settable options and avoid 
automatically forwarding them to execution workers.
   
   ### Fresh table resolution before planning
   
   - Load referenced tables through the schema provider for each query.
   - Build query-local table providers so schema resolution and scan planning 
use the same loaded metadata.
   - Let subsequent queries observe updates without rebuilding the shared 
catalog client.
   
   ### Reusable providers with replaceable data sources
   
   DataFusion’s DataSourceExec is a standard physical scan node that delegates 
reading, partitioning, statistics, and metrics to a DataSource implementation. 
Keeping scan construction replaceable would let other engines (like 
dataFusion-distributed) reuse the upstream catalog, schema, and table providers 
while supplying its own data source; the execution details belong in a separate 
issue.
   
   ## Proposed architecture
   
   ```text
   Embedding application
     └── authenticates request and supplies trusted context
             |
             v
   Query-bound catalog/schema resolution
     ├── uses the shared catalog client with that context
     ├── loads referenced tables for this query
     └── creates query-local table providers
             |
             v
   Iceberg TableProvider
     ├── schema(): exposes the loaded table's schema
     └── scan():
           ├── selects snapshot
           ├── converts predicates
           ├── invokes Iceberg's scan planner
           └── passes planned FileScanTasks to the factory
             |
             v
   Configurable data-source factory
     └── organizes planned tasks for local or distributed execution
             |
             v
   DataSourceExec
     ├── local IcebergDataSource
     └── downstream/distributed IcebergDataSource
             |
             v
   Execution
     └── uses Iceberg's Arrow reader to read assigned tasks
   ```
   
   I can go more into detail about how I see these structs in a subsequent 
issue, or edit and append here.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to