Hi all, a question about GlueCatalog and what it writes to the Glue table on 
commit.

GlueTableOperations.prepareProperties<https://github.com/apache/iceberg/blob/d008ad230bed30383a667af8151573e48de7624c/aws/src/main/java/org/apache/iceberg/aws/glue/GlueTableOperations.java#L291-L301>
 sets only table_type, metadata_location and previous_metadata_locationin the 
Glue table Parameters. It carries over whatever parameters the Glue table 
already had, but never copies TableMetadata.properties().

The Hive catalog works differently. 
HMSTablePropertyHelper.updateHmsTableForIcebergTable<https://github.com/apache/iceberg/blob/d008ad230bed30383a667af8151573e48de7624c/hive-metastore/src/main/java/org/apache/iceberg/hive/HMSTablePropertyHelper.java#L78-L130>
 pushes all Iceberg table properties into HMS params, plus the current snapshot 
summary, schema, partition spec and sort order, all capped by 
iceberg.hive.table-property-max-size 
(https://github.com/apache/iceberg/blob/d008ad230bed30383a667af8151573e48de7624c/hive-metastore/src/main/java/org/apache/iceberg/hive/HiveOperationsBase.java#L53).
 So the same table shows a lot more metadata through HMS than through Glue.

Questions:

  1.  Is the minimal Glue footprint intentional? The javadoc says Glue info is 
"informational only" and metadata_location is the source of truth. Or has 
nobody needed more yet?
  2.  Would the community accept a change that mirrors the Hive behavior for 
Glue? That would mean syncing table properties (and maybe the snapshot 
summary/schema/spec) into Glue Parameters, with a size cap and an opt-in flag. 
Open design questions would be Glue's parameter size limits, removing keys that 
were deleted in Iceberg, and protecting reserved keys.


Our use case: tools that only read the Glue catalog (governance, discovery, 
Lake Formation-based tooling) can't see Iceberg table properties without 
opening the metadata JSON.

Slack message - 
https://apache-iceberg.slack.com/archives/C025PH0G1D4/p1790773851197099


Reply via email to