Vidit Gupta created HIVE-30096:
----------------------------------
Summary: HiveAlterHandler forces a full file listing to recompute
stats on every table-level alter
Key: HIVE-30096
URL: https://issues.apache.org/jira/browse/HIVE-30096
Project: Hive
Issue Type: Bug
Components: Standalone Metastore
Affects Versions: 4.2.0
Reporter: Vidit Gupta
Assignee: Vidit Gupta
(final version, includes the legacy-table finding and the trade-off note):
For any non-rename, table-level alter of an unpartitioned table,
HiveAlterHandler.alterTable calls MetaStoreServerUtils.updateTableStatsSlow
with forceRecompute=true unconditionally. That makes the existing shortcut that
reuses the fast stats already present in the table parameters (the
!forceRecompute && containsAllFastStats check) unreachable, so the metastore
recursively lists the entire table location and materializes one FileStatus per
file in memory, inside a single RPC.
Nothing on the server gates this path: requireCalStats(null, null, tbl, ec)
returns true unconditionally for table-level alters (both partition arguments
are null), metastore.stats.autogather is only consulted on the
create/add-partition paths (HMSHandler.canUpdateStats), the DO_NOT_UPDATE_STATS
table parameter is transient (removed on first use, see the HIVE-10228 note in
updateTableStatsSlow), and DO_NOT_UPDATE_STATS in the EnvironmentContext
requires changing every client.
This is catastrophic for table formats that keep partitioning in their own
transaction log (Delta, Iceberg, Hudi): they register as unpartitioned tables
(getPartitionKeysSize() == 0) with millions of files under one directory. On
our deployment, a property-only alter of a Delta table with ~4M files on GCS
materialized ~7.9GB of FileStatus objects and OOMed the metastore twice with an
8GB heap. Sending DO_NOT_UPDATE_STATS in the EnvironmentContext from the client
avoids it (the same alter then applies instantly at flat heap), confirming the
listing is the cause. The recomputed values are also of little use for such
tables: a directory listing counts files that are no longer part of the current
snapshot (e.g. not yet vacuumed), and for any non-ANALYZE alter the recomputed
stats are immediately marked not accurate by setBasicStatsState(FALSE) anyway.
Proposed fix: force the recompute only when the alter actually changed the
table location. Metadata-only alters then take the existing
containsAllFastStats shortcut; location changes recompute exactly as before;
tables missing the fast stats are still listed once and self-heal. This also
makes the alter path consistent with the create path, which already passes
forceRecompute=false.
Two notes: (1) tables created by Hive 2.x lack the numFilesErasureCoded
parameter (added in 3.0), so their first alter still lists once — operators
migrating old warehouses can backfill that key (value 0 where erasure coding
does not apply) to avoid even one listing on very large tables. (2) After this
change, refreshing stale quick stats without a location change requires ANALYZE
TABLE (or happens via StatsTask on Hive-managed writes), rather than happening
as a side effect of any alter. PR to follow.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)