[
https://issues.apache.org/jira/browse/HIVE-30096?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated HIVE-30096:
----------------------------------
Labels: pull-request-available (was: )
> HiveAlterHandler forces a full file listing to recompute stats on every
> table-level alter
> -----------------------------------------------------------------------------------------
>
> Key: HIVE-30096
> URL: https://issues.apache.org/jira/browse/HIVE-30096
> Project: Hive
> Issue Type: Bug
> Components: Standalone Metastore
> Affects Versions: 4.2.0
> Reporter: Vidit Gupta
> Assignee: Vidit Gupta
> Priority: Major
> Labels: pull-request-available
>
> (final version, includes the legacy-table finding and the trade-off note):
> For any non-rename, table-level alter of an unpartitioned table,
> HiveAlterHandler.alterTable calls MetaStoreServerUtils.updateTableStatsSlow
> with forceRecompute=true unconditionally. That makes the existing shortcut
> that reuses the fast stats already present in the table parameters (the
> !forceRecompute && containsAllFastStats check) unreachable, so the metastore
> recursively lists the entire table location and materializes one FileStatus
> per file in memory, inside a single RPC.
> Nothing on the server gates this path: requireCalStats(null, null, tbl, ec)
> returns true unconditionally for table-level alters (both partition arguments
> are null), metastore.stats.autogather is only consulted on the
> create/add-partition paths (HMSHandler.canUpdateStats), the
> DO_NOT_UPDATE_STATS table parameter is transient (removed on first use, see
> the HIVE-10228 note in updateTableStatsSlow), and DO_NOT_UPDATE_STATS in the
> EnvironmentContext requires changing every client.
> This is catastrophic for table formats that keep partitioning in their own
> transaction log (Delta, Iceberg, Hudi): they register as unpartitioned tables
> (getPartitionKeysSize() == 0) with millions of files under one directory. On
> our deployment, a property-only alter of a Delta table with ~4M files on GCS
> materialized ~7.9GB of FileStatus objects and OOMed the metastore twice with
> an 8GB heap. Sending DO_NOT_UPDATE_STATS in the EnvironmentContext from the
> client avoids it (the same alter then applies instantly at flat heap),
> confirming the listing is the cause. The recomputed values are also of little
> use for such tables: a directory listing counts files that are no longer part
> of the current snapshot (e.g. not yet vacuumed), and for any non-ANALYZE
> alter the recomputed stats are immediately marked not accurate by
> setBasicStatsState(FALSE) anyway.
> Proposed fix: force the recompute only when the alter actually changed the
> table location. Metadata-only alters then take the existing
> containsAllFastStats shortcut; location changes recompute exactly as before;
> tables missing the fast stats are still listed once and self-heal. This also
> makes the alter path consistent with the create path, which already passes
> forceRecompute=false.
> Two notes: (1) tables created by Hive 2.x lack the numFilesErasureCoded
> parameter (added in 3.0), so their first alter still lists once — operators
> migrating old warehouses can backfill that key (value 0 where erasure coding
> does not apply) to avoid even one listing on very large tables. (2) After
> this change, refreshing stale quick stats without a location change requires
> ANALYZE TABLE (or happens via StatsTask on Hive-managed writes), rather than
> happening as a side effect of any alter. PR to follow.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)