Vidit Gupta created HIVE-30096:
----------------------------------

             Summary: HiveAlterHandler forces a full file listing to recompute 
stats on every table-level alter
                 Key: HIVE-30096
                 URL: https://issues.apache.org/jira/browse/HIVE-30096
             Project: Hive
          Issue Type: Bug
          Components: Standalone Metastore
    Affects Versions: 4.2.0
            Reporter: Vidit Gupta
            Assignee: Vidit Gupta


(final version, includes the legacy-table finding and the trade-off note):

For any non-rename, table-level alter of an unpartitioned table, 
HiveAlterHandler.alterTable calls MetaStoreServerUtils.updateTableStatsSlow 
with forceRecompute=true unconditionally. That makes the existing shortcut that 
reuses the fast stats already present in the table parameters (the 
!forceRecompute && containsAllFastStats check) unreachable, so the metastore 
recursively lists the entire table location and materializes one FileStatus per 
file in memory, inside a single RPC.

Nothing on the server gates this path: requireCalStats(null, null, tbl, ec) 
returns true unconditionally for table-level alters (both partition arguments 
are null), metastore.stats.autogather is only consulted on the 
create/add-partition paths (HMSHandler.canUpdateStats), the DO_NOT_UPDATE_STATS 
table parameter is transient (removed on first use, see the HIVE-10228 note in 
updateTableStatsSlow), and DO_NOT_UPDATE_STATS in the EnvironmentContext 
requires changing every client.

This is catastrophic for table formats that keep partitioning in their own 
transaction log (Delta, Iceberg, Hudi): they register as unpartitioned tables 
(getPartitionKeysSize() == 0) with millions of files under one directory. On 
our deployment, a property-only alter of a Delta table with ~4M files on GCS 
materialized ~7.9GB of FileStatus objects and OOMed the metastore twice with an 
8GB heap. Sending DO_NOT_UPDATE_STATS in the EnvironmentContext from the client 
avoids it (the same alter then applies instantly at flat heap), confirming the 
listing is the cause. The recomputed values are also of little use for such 
tables: a directory listing counts files that are no longer part of the current 
snapshot (e.g. not yet vacuumed), and for any non-ANALYZE alter the recomputed 
stats are immediately marked not accurate by setBasicStatsState(FALSE) anyway.

Proposed fix: force the recompute only when the alter actually changed the 
table location. Metadata-only alters then take the existing 
containsAllFastStats shortcut; location changes recompute exactly as before; 
tables missing the fast stats are still listed once and self-heal. This also 
makes the alter path consistent with the create path, which already passes 
forceRecompute=false.

Two notes: (1) tables created by Hive 2.x lack the numFilesErasureCoded 
parameter (added in 3.0), so their first alter still lists once — operators 
migrating old warehouses can backfill that key (value 0 where erasure coding 
does not apply) to avoid even one listing on very large tables. (2) After this 
change, refreshing stale quick stats without a location change requires ANALYZE 
TABLE (or happens via StatsTask on Hive-managed writes), rather than happening 
as a side effect of any alter. PR to follow.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to