[
https://issues.apache.org/jira/browse/HIVE-29843?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Zhenqi Li updated HIVE-29843:
-----------------------------
Description:
When multiple clients concurrently initialize the same Iceberg namespace, Hive
Metastore may process identical {{create_database}} requests and create the
same database directory. Only one request can successfully insert the database
metadata; the other requests fail because of the database unique constraint.
The failed requests then use a local {{madeDir}} or {{madeExternalDir}} flag to
decide whether to recursively clean up the entire database directory. However,
this flag only records that the request created or observed a usable directory;
it does not prove that the request still has exclusive ownership of that
directory. Another request may already have committed metadata for the same
path or written data below it.
As a result, a failed {{create_database}} request may move or delete a shared
database directory, leaving valid Metastore metadata pointing to missing
object-storage files.
h2. Affected Versions
The issue was reproduced against Hive Metastore {{{}4.0.0-beta-1{}}}.
Source inspection found the same directory-ownership and rollback-cleanup
pattern in Hive branches from {{1.0}} through {{{}4.2.1{}}}. I did not find an
upstream fix specifically addressing the {{create_database}} directory
ownership race.
h2. Description
Multiple clients initializing the same Iceberg namespace can concurrently call
Hive Metastore with identical {{create_database}} requests.
A typical target path is:
{code:java}
s3a://warehouse/iceberg/test1.db
{code}
The following sequence can occur:
# Multiple requests check that the database does not exist.
# Multiple requests check that the target directory does not exist.
# Multiple requests call {{mkdirs}} for the same path.
# One or more requests record {{madeDir}} or {{madeExternalDir}} as {{true}}
before the metadata winner is determined.
# One request successfully inserts the database into the Metastore {{DBS}}
table.
# The other requests fail because of the {{DBS.UNIQUE_DATABASE}} unique
constraint.
# The failed requests enter the rollback cleanup path.
# Based on their local flag, they assume that they own the database directory.
# A failed request recursively moves or deletes the entire database directory.
There is also a more general interleaving, similar to the one documented in
HIVE-23437:
# Request R1 creates the directory and is paused before committing metadata.
# Request R2 observes that the directory already exists, so R2 does not create
it.
# R2 successfully commits database metadata referencing that directory.
# R1 resumes, loses the metadata race, and rolls back.
# R1 still has {{madeDir=true}} and recursively deletes the directory already
adopted by R2.
This second interleaving shows that even a perfectly atomic {{mkdir}}
implementation would not be sufficient. Directory creation provenance is not
equivalent to current deletion ownership.
h2. Steps to Reproduce
h3. 1. Prepare Hive Metastore
Start Hive Metastore with a MySQL metadata database and S3A-compatible object
storage. Confirm that the target database path does not exist.
For example:
{code:java}
s3a://warehouse/hms_direct_race.db
{code}
h3. 2. Concurrently invoke create_database
Use multiple independent Hive Metastore clients and concurrently call:
{code:java}
client.createDatabase(database);
{code}
All requests must use the same catalog, database name, and location. The
standalone reproduction used:
{code:java}
catalog: hive
database: hms_direct_race
concurrency: 16
{code}
h3. 3. Observe the Metastore result
The reproduced result was:
* 1 request succeeded.
* 15 requests failed.
* The failed requests reported a MySQL unique-key conflict:
{code:java}
Duplicate entry 'hms_direct_race-hive'
for key 'DBS.UNIQUE_DATABASE'
{code}
* Multiple handlers logged that the same database path was created.
* Failed requests entered the database-directory cleanup path.
h3. 4. Observe directory cleanup
If another request writes content below the database directory before rollback
cleanup runs, the failed request may move or delete that content together with
the entire database directory.
For example, affected Iceberg content may include:
{code:java}
iceberg/test1.db/test3/metadata/00000-....metadata.json
{code}
When Hadoop Trash is enabled, the path may be moved under a location similar to:
{code:java}
s3a://warehouse/user/root/.Trash/Current/iceberg/test1.db
{code}
h2. Reproduction Log Excerpts
The following logs are from a standalone reproduction that bypassed all
higher-level application code. It used 16 independent Hive Metastore clients.
The run started at {{2026-08-27 02:23:40 UTC}} and targeted database
{{{}hms_direct_race_20260827T022340Z{}}}.
h3. Client result
{code:java}
HMS_CREATE_DATABASE_RACE host=hive-metastore port=9083
database=hms_direct_race_20260827T022340Z workers=16
location=s3a://warehouse/hms_direct_race_20260827T022340Z.db
worker=00 result=FAILURE error=MetaException: JDODataStoreException: Exception
thrown flushing changes to datastore Root cause:
java.sql.SQLIntegrityConstraintViolationException: Duplicate entry
'hms_direct_race_20260827t022340z-hive' for key 'DBS.UNIQUE_DATABASE'
worker=01 result=FAILURE error=MetaException: JDODataStoreException: Exception
thrown flushing changes to datastore Root cause:
java.sql.SQLIntegrityConstraintViolationException: Duplicate entry
'hms_direct_race_20260827t022340z-hive' for key 'DBS.UNIQUE_DATABASE'
...
worker=11 result=SUCCESS
...
worker=15 result=FAILURE error=MetaException: JDODataStoreException: Exception
thrown flushing changes to datastore Root cause:
java.sql.SQLIntegrityConstraintViolationException: Duplicate entry
'hms_direct_race_20260827t022340z-hive' for key 'DBS.UNIQUE_DATABASE'
HMS_CREATE_DATABASE_RACE_COMPLETE elapsed_ms=10380
{code}
The final result was one successful request and 15 failures caused by
{{{}DBS.UNIQUE_DATABASE{}}}.
h3. HMS audit and directory creation
Hive Metastore audit logs recorded multiple concurrent {{create_database}}
requests for the same database. One example is:
{code:java}
2026-08-27T02:23:48,142 INFO [Metastore-Handler-Pool: Thread-1341]
HiveMetaStore.audit: ugi=root ip=172.19.0.6 cmd=source:172.19.0.6
create_database: Database(name:hms_direct_race_20260827T022340Z,
description:direct concurrent create_database test,
locationUri:s3a://warehouse/hms_direct_race_20260827T022340Z.db, parameters:{},
catalogName:hive)
{code}
Multiple handlers then processed the same database path:
{code:java}
2026-08-27T02:23:48,401 INFO [Metastore-Handler-Pool: Thread-1340] HMSHandler:
Creating database path in external directory
s3a://warehouse/hms_direct_race_20260827T022340Z.db
2026-08-27T02:23:48,403 INFO [Metastore-Handler-Pool: Thread-1334] HMSHandler:
Creating database path in external directory
s3a://warehouse/hms_direct_race_20260827T022340Z.db
2026-08-27T02:23:48,434 INFO [Metastore-Handler-Pool: Thread-1338] HMSHandler:
Created database path in external directory
s3a://warehouse/hms_direct_race_20260827T022340Z.db
2026-08-27T02:23:48,438 INFO [Metastore-Handler-Pool: Thread-1334] HMSHandler:
Created database path in external directory
s3a://warehouse/hms_direct_race_20260827T022340Z.db
{code}
h3. Sentinel and cleanup
A watcher observed the database directory during the race window and uploaded a
sentinel object below it:
{code:java}
WATCHER_START 2026-08-27T02:23:41Z
DIRECTORY_VISIBLE 2026-08-27T02:23:48Z
s3://warehouse/hms_direct_race_20260827T022340Z.db
46 bytes ->
`minio/warehouse/hms_direct_race_20260827T022340Z.db/sentinel/metadata/race-sentinel.txt`
SENTINEL_UPLOAD_DONE 2026-08-27T02:23:48Z
WATCHER_END 2026-08-27T02:23:48Z
{code}
The unique-key conflict and subsequent cleanup appeared together in the Hive
Metastore logs:
{code:java}
Caused by: java.sql.SQLIntegrityConstraintViolationException: Duplicate entry
'hms_direct_race_20260827t022340z-hive' for key 'DBS.UNIQUE_DATABASE'
at
org.apache.hadoop.hive.metastore.HMSHandler.create_database_core(HMSHandler.java:1274)
at
org.apache.hadoop.hive.metastore.HMSHandler.create_database(HMSHandler.java:1343)
2026-08-27T02:23:58,482 WARN [Metastore-Handler-Pool: Thread-1342] FileUtils:
No such file or directory: s3a://warehouse/hms_direct_race_20260827T022340Z.db;
Force to delete it.
2026-08-27T02:23:58,487 ERROR [Metastore-Handler-Pool: Thread-1341] FileUtils:
Failed to delete s3a://warehouse/hms_direct_race_20260827T022340Z.db
{code}
After the test, no objects remained under the original path or the
corresponding Trash path:
{code:java}
=== OBJECTS UNDER DIRECT PATH ===
=== OBJECTS UNDER TRASH PATH ===
=== DIRECT PATH STAT ===
mc: <ERROR> Unable to stat
`minio/warehouse/hms_direct_race_20260827T022340Z.db/`. Object does not exist.
{code}
At the same time, the database row still existed in the MySQL {{DBS}} table:
{code:java}
DB_ID NAME CTLG_NAME DB_LOCATION_URI
104 hms_direct_race_20260827t022340z hive
s3a://warehouse/hms_direct_race_20260827T022340Z.db
{code}
Trash was disabled in this standalone reproduction, so cleanup directly deleted
the directory instead of moving it to {{{}.Trash/Current{}}}.
h2. Actual Result
A failed {{create_database}} request may recursively clean up the entire final
database directory rather than cleaning up only request-private resources.
This can leave the system in the following state:
* Database metadata still exists in Hive Metastore.
* Table metadata may still exist in Hive Metastore.
* Iceberg {{metadata_location}} values still reference the original
object-storage paths.
* The corresponding Iceberg metadata files have been moved to Trash or
recursively deleted.
* Subsequent Iceberg queries, startup recovery, or metadata commits receive
404/not-found errors.
* Iceberg tables become unreadable or unwritable.
h2. Expected Result
When multiple requests concurrently create the same database:
* One request should successfully create the database.
* Other requests should receive {{AlreadyExistsException}} or an equivalent
database-already-exists error.
* Failed requests must not delete, move, or otherwise modify the shared final
database directory.
* Failed requests may only recursively clean up resources that are provably
private to that request, such as a request-scoped staging path.
h2. Root Cause
{{create_database}} uses a non-atomic check-then-create sequence:
{code:java}
get_database_core(...);
if (database does not exist) {
create_database_core(...);
}
{code}
Inside {{{}create_database_core{}}}, filesystem ownership is inferred from
logic similar to:
{code:java}
if (!wh.isDir(dbPath)) {
if (!wh.mkdirs(dbPath)) {
throw new MetaException(...);
}
madeDir = true;
}
{code}
A successful {{mkdirs}} call only shows that the directory became available
through this request. It does not prove continuing exclusive ownership.
Multiple requests may set {{madeDir}} to {{{}true{}}}, and even a directory
created exclusively by one request may subsequently be adopted by a different
request that wins the metadata transaction.
When one request commits the Metastore metadata and other requests fail because
of {{{}DBS.UNIQUE_DATABASE{}}}, the failed requests execute cleanup similar to:
{code:java}
if (!success) {
ms.rollbackTransaction();
if (madeDir) {
wh.deleteDir(dbPath, true);
}
}
{code}
The corresponding external database path cleanup in Hive {{4.0.0-beta-1}} is
equivalent to:
{code:java}
if (madeExternalDir && dbExtPath != null) {
wh.deleteDir(dbExtPath, true, db);
}
{code}
The cleanup scope is the entire final database directory, not an empty
request-private directory or an object written only by the failed request.
The unique-key conflict itself is an expected concurrency outcome. The
data-loss risk comes from treating local directory-creation provenance as
authorization to recursively delete a shared final location after metadata
rollback.
h2. Source References
* Hive {{4.0.0-beta-1}} {{{}HMSHandler.create_database_core{}}}:
[source|https://github.com/apache/hive/blob/rel/release-4.0.0-beta-1/standalone-metastore/metastore-server/src/main/java/org/apache/hadoop/hive/metastore/HMSHandler.java#L1138-L1341]
* Hive {{3.1.2}} {{{}create_database_core{}}}:
[source|https://github.com/apache/hive/blob/rel/release-3.1.2/standalone-metastore/src/main/java/org/apache/hadoop/hive/metastore/HiveMetaStore.java#L1240-L1287]
* Hive {{4.2.1}} {{{}HMSHandler.create_database_core{}}}:
[source|https://github.com/apache/hive/blob/rel/release-4.2.1/standalone-metastore/metastore-server/src/main/java/org/apache/hadoop/hive/metastore/HMSHandler.java#L1163-L1347]
was:
When multiple clients concurrently initialize the same Iceberg namespace, Hive
Metastore may process identical {{create_database}} requests and create the
same database directory. Only one request can successfully insert the database
metadata; the other requests fail because of the database unique constraint.
The failed requests then use a local {{madeDir}} or {{madeExternalDir}} flag to
decide whether to recursively clean up the entire database directory. However,
this flag only records that the request created or observed a usable directory;
it does not prove that the request still has exclusive ownership of that
directory. Another request may already have committed metadata for the same
path or written data below it.
As a result, a failed {{create_database}} request may move or delete a shared
database directory, leaving valid Metastore metadata pointing to missing
object-storage files.
h2. Affected Versions
The issue was reproduced against Hive Metastore {{{}4.0.0-beta-1{}}}.
Source inspection found the same directory-ownership and rollback-cleanup
pattern in Hive branches from {{1.0}} through {{{}4.2.1{}}}. I did not find an
upstream fix specifically addressing the {{create_database}} directory
ownership race.
h2. Description
Multiple clients initializing the same Iceberg namespace can concurrently call
Hive Metastore with identical {{create_database}} requests.
A typical target path is:
{code:java}
s3a://warehouse/iceberg/test1.db
{code}
The following sequence can occur:
# Multiple requests check that the database does not exist.
# Multiple requests check that the target directory does not exist.
# Multiple requests call {{mkdirs}} for the same path.
# Multiple requests set a local {{madeDir}} or {{madeExternalDir}} flag to
{{{}true{}}}.
# One request successfully inserts the database into the Metastore {{DBS}}
table.
# The other requests fail because of the {{DBS.UNIQUE_DATABASE}} unique
constraint.
# The failed requests enter the rollback cleanup path.
# Based on their local flag, they assume that they own the database directory.
# A failed request recursively moves or deletes the entire database directory.
There is also a more general interleaving, similar to the one documented in
HIVE-23437:
# Request R1 creates the directory and is paused before committing metadata.
# Request R2 observes that the directory already exists, so R2 does not create
it.
# R2 successfully commits database metadata referencing that directory.
# R1 resumes, loses the metadata race, and rolls back.
# R1 still has {{madeDir=true}} and recursively deletes the directory already
adopted by R2.
This second interleaving shows that even a perfectly atomic {{mkdir}}
implementation would not be sufficient. Directory creation provenance is not
equivalent to current deletion ownership.
h2. Steps to Reproduce
h3. 1. Prepare Hive Metastore
Start Hive Metastore with a MySQL metadata database and S3A-compatible object
storage. Confirm that the target database path does not exist.
For example:
{code:java}
s3a://warehouse/hms_direct_race.db
{code}
h3. 2. Concurrently invoke create_database
Use multiple independent Hive Metastore clients and concurrently call:
{code:java}
client.createDatabase(database);
{code}
All requests must use the same catalog, database name, and location. The
standalone reproduction used:
{code:java}
catalog: hive
database: hms_direct_race
concurrency: 16
{code}
h3. 3. Observe the Metastore result
The reproduced result was:
* 1 request succeeded.
* 15 requests failed.
* The failed requests reported a MySQL unique-key conflict:
{code:java}
Duplicate entry 'hms_direct_race-hive'
for key 'DBS.UNIQUE_DATABASE'
{code}
* Multiple handlers logged that the same database path was created.
* Failed requests entered the database-directory cleanup path.
h3. 4. Observe directory cleanup
If another request writes content below the database directory before rollback
cleanup runs, the failed request may move or delete that content together with
the entire database directory.
For example, affected Iceberg content may include:
{code:java}
iceberg/test1.db/test3/metadata/00000-....metadata.json
{code}
When Hadoop Trash is enabled, the path may be moved under a location similar to:
{code:java}
s3a://warehouse/user/root/.Trash/Current/iceberg/test1.db
{code}
h2. Reproduction Log Excerpts
The following logs are from a standalone reproduction that bypassed all
higher-level application code. It used 16 independent Hive Metastore clients.
The run started at {{2026-08-27 02:23:40 UTC}} and targeted database
{{{}hms_direct_race_20260827T022340Z{}}}.
h3. Client result
{code:java}
HMS_CREATE_DATABASE_RACE host=hive-metastore port=9083
database=hms_direct_race_20260827T022340Z workers=16
location=s3a://warehouse/hms_direct_race_20260827T022340Z.db
worker=00 result=FAILURE error=MetaException: JDODataStoreException: Exception
thrown flushing changes to datastore Root cause:
java.sql.SQLIntegrityConstraintViolationException: Duplicate entry
'hms_direct_race_20260827t022340z-hive' for key 'DBS.UNIQUE_DATABASE'
worker=01 result=FAILURE error=MetaException: JDODataStoreException: Exception
thrown flushing changes to datastore Root cause:
java.sql.SQLIntegrityConstraintViolationException: Duplicate entry
'hms_direct_race_20260827t022340z-hive' for key 'DBS.UNIQUE_DATABASE'
...
worker=11 result=SUCCESS
...
worker=15 result=FAILURE error=MetaException: JDODataStoreException: Exception
thrown flushing changes to datastore Root cause:
java.sql.SQLIntegrityConstraintViolationException: Duplicate entry
'hms_direct_race_20260827t022340z-hive' for key 'DBS.UNIQUE_DATABASE'
HMS_CREATE_DATABASE_RACE_COMPLETE elapsed_ms=10380
{code}
The final result was one successful request and 15 failures caused by
{{{}DBS.UNIQUE_DATABASE{}}}.
h3. HMS audit and directory creation
Hive Metastore audit logs recorded multiple concurrent {{create_database}}
requests for the same database. One example is:
{code:java}
2026-08-27T02:23:48,142 INFO [Metastore-Handler-Pool: Thread-1341]
HiveMetaStore.audit: ugi=root ip=172.19.0.6 cmd=source:172.19.0.6
create_database: Database(name:hms_direct_race_20260827T022340Z,
description:direct concurrent create_database test,
locationUri:s3a://warehouse/hms_direct_race_20260827T022340Z.db, parameters:{},
catalogName:hive)
{code}
Multiple handlers then processed the same database path:
{code:java}
2026-08-27T02:23:48,401 INFO [Metastore-Handler-Pool: Thread-1340] HMSHandler:
Creating database path in external directory
s3a://warehouse/hms_direct_race_20260827T022340Z.db
2026-08-27T02:23:48,403 INFO [Metastore-Handler-Pool: Thread-1334] HMSHandler:
Creating database path in external directory
s3a://warehouse/hms_direct_race_20260827T022340Z.db
2026-08-27T02:23:48,434 INFO [Metastore-Handler-Pool: Thread-1338] HMSHandler:
Created database path in external directory
s3a://warehouse/hms_direct_race_20260827T022340Z.db
2026-08-27T02:23:48,438 INFO [Metastore-Handler-Pool: Thread-1334] HMSHandler:
Created database path in external directory
s3a://warehouse/hms_direct_race_20260827T022340Z.db
{code}
h3. Sentinel and cleanup
A watcher observed the database directory during the race window and uploaded a
sentinel object below it:
{code:java}
WATCHER_START 2026-08-27T02:23:41Z
DIRECTORY_VISIBLE 2026-08-27T02:23:48Z
s3://warehouse/hms_direct_race_20260827T022340Z.db
46 bytes ->
`minio/warehouse/hms_direct_race_20260827T022340Z.db/sentinel/metadata/race-sentinel.txt`
SENTINEL_UPLOAD_DONE 2026-08-27T02:23:48Z
WATCHER_END 2026-08-27T02:23:48Z
{code}
The unique-key conflict and subsequent cleanup appeared together in the Hive
Metastore logs:
{code:java}
Caused by: java.sql.SQLIntegrityConstraintViolationException: Duplicate entry
'hms_direct_race_20260827t022340z-hive' for key 'DBS.UNIQUE_DATABASE'
at
org.apache.hadoop.hive.metastore.HMSHandler.create_database_core(HMSHandler.java:1274)
at
org.apache.hadoop.hive.metastore.HMSHandler.create_database(HMSHandler.java:1343)
2026-08-27T02:23:58,482 WARN [Metastore-Handler-Pool: Thread-1342] FileUtils:
No such file or directory: s3a://warehouse/hms_direct_race_20260827T022340Z.db;
Force to delete it.
2026-08-27T02:23:58,487 ERROR [Metastore-Handler-Pool: Thread-1341] FileUtils:
Failed to delete s3a://warehouse/hms_direct_race_20260827T022340Z.db
{code}
After the test, no objects remained under the original path or the
corresponding Trash path:
{code:java}
=== OBJECTS UNDER DIRECT PATH ===
=== OBJECTS UNDER TRASH PATH ===
=== DIRECT PATH STAT ===
mc: <ERROR> Unable to stat
`minio/warehouse/hms_direct_race_20260827T022340Z.db/`. Object does not exist.
{code}
At the same time, the database row still existed in the MySQL {{DBS}} table:
{code:java}
DB_ID NAME CTLG_NAME DB_LOCATION_URI
104 hms_direct_race_20260827t022340z hive
s3a://warehouse/hms_direct_race_20260827T022340Z.db
{code}
Trash was disabled in this standalone reproduction, so cleanup directly deleted
the directory instead of moving it to {{{}.Trash/Current{}}}.
h2. Actual Result
A failed {{create_database}} request may recursively clean up the entire final
database directory rather than cleaning up only request-private resources.
This can leave the system in the following state:
* Database metadata still exists in Hive Metastore.
* Table metadata may still exist in Hive Metastore.
* Iceberg {{metadata_location}} values still reference the original
object-storage paths.
* The corresponding Iceberg metadata files have been moved to Trash or
recursively deleted.
* Subsequent Iceberg queries, startup recovery, or metadata commits receive
404/not-found errors.
* Iceberg tables become unreadable or unwritable.
h2. Expected Result
When multiple requests concurrently create the same database:
* One request should successfully create the database.
* Other requests should receive {{AlreadyExistsException}} or an equivalent
database-already-exists error.
* Failed requests must not delete, move, or otherwise modify the shared final
database directory.
* Failed requests may only recursively clean up resources that are provably
private to that request, such as a request-scoped staging path.
h2. Root Cause
{{create_database}} uses a non-atomic check-then-create sequence:
{code:java}
get_database_core(...);
if (database does not exist) {
create_database_core(...);
}
{code}
Inside {{{}create_database_core{}}}, filesystem ownership is inferred from
logic similar to:
{code:java}
if (!wh.isDir(dbPath)) {
if (!wh.mkdirs(dbPath)) {
throw new MetaException(...);
}
madeDir = true;
}
{code}
A successful {{mkdirs}} call only shows that the directory became available
through this request. It does not prove continuing exclusive ownership.
Multiple requests may set {{madeDir}} to {{{}true{}}}, and even a directory
created exclusively by one request may subsequently be adopted by a different
request that wins the metadata transaction.
When one request commits the Metastore metadata and other requests fail because
of {{{}DBS.UNIQUE_DATABASE{}}}, the failed requests execute cleanup similar to:
{code:java}
if (!success) {
ms.rollbackTransaction();
if (madeDir) {
wh.deleteDir(dbPath, true);
}
}
{code}
The corresponding external database path cleanup in Hive {{4.0.0-beta-1}} is
equivalent to:
{code:java}
if (madeExternalDir && dbExtPath != null) {
wh.deleteDir(dbExtPath, true, db);
}
{code}
The cleanup scope is the entire final database directory, not an empty
request-private directory or an object written only by the failed request.
The unique-key conflict itself is an expected concurrency outcome. The
data-loss risk comes from treating local directory-creation provenance as
authorization to recursively delete a shared final location after metadata
rollback.
> Concurrent create_database requests may recursively delete a shared database
> directory
> --------------------------------------------------------------------------------------
>
> Key: HIVE-29843
> URL: https://issues.apache.org/jira/browse/HIVE-29843
> Project: Hive
> Issue Type: Bug
> Components: Standalone Metastore
> Affects Versions: 4.0.0-beta-1
> Environment: Hive Metastore: 4.0.0-beta-1
> MySQL: 8.0.30
> Hadoop: 3.3.1
> FileSystem: S3A-compatible object storage
> Standalone reproduction object storage: MinIO
> Iceberg: 1.6.1
> Concurrency: 16 independent Hive Metastore clients
> Reporter: Zhenqi Li
> Assignee: Zhenqi Li
> Priority: Major
>
> When multiple clients concurrently initialize the same Iceberg namespace,
> Hive Metastore may process identical {{create_database}} requests and create
> the same database directory. Only one request can successfully insert the
> database metadata; the other requests fail because of the database unique
> constraint.
> The failed requests then use a local {{madeDir}} or {{madeExternalDir}} flag
> to decide whether to recursively clean up the entire database directory.
> However, this flag only records that the request created or observed a usable
> directory; it does not prove that the request still has exclusive ownership
> of that directory. Another request may already have committed metadata for
> the same path or written data below it.
> As a result, a failed {{create_database}} request may move or delete a shared
> database directory, leaving valid Metastore metadata pointing to missing
> object-storage files.
> h2. Affected Versions
> The issue was reproduced against Hive Metastore {{{}4.0.0-beta-1{}}}.
> Source inspection found the same directory-ownership and rollback-cleanup
> pattern in Hive branches from {{1.0}} through {{{}4.2.1{}}}. I did not find
> an upstream fix specifically addressing the {{create_database}} directory
> ownership race.
> h2. Description
> Multiple clients initializing the same Iceberg namespace can concurrently
> call Hive Metastore with identical {{create_database}} requests.
> A typical target path is:
> {code:java}
> s3a://warehouse/iceberg/test1.db
> {code}
> The following sequence can occur:
> # Multiple requests check that the database does not exist.
> # Multiple requests check that the target directory does not exist.
> # Multiple requests call {{mkdirs}} for the same path.
> # One or more requests record {{madeDir}} or {{madeExternalDir}} as {{true}}
> before the metadata winner is determined.
> # One request successfully inserts the database into the Metastore {{DBS}}
> table.
> # The other requests fail because of the {{DBS.UNIQUE_DATABASE}} unique
> constraint.
> # The failed requests enter the rollback cleanup path.
> # Based on their local flag, they assume that they own the database
> directory.
> # A failed request recursively moves or deletes the entire database
> directory.
> There is also a more general interleaving, similar to the one documented in
> HIVE-23437:
> # Request R1 creates the directory and is paused before committing metadata.
> # Request R2 observes that the directory already exists, so R2 does not
> create it.
> # R2 successfully commits database metadata referencing that directory.
> # R1 resumes, loses the metadata race, and rolls back.
> # R1 still has {{madeDir=true}} and recursively deletes the directory
> already adopted by R2.
> This second interleaving shows that even a perfectly atomic {{mkdir}}
> implementation would not be sufficient. Directory creation provenance is not
> equivalent to current deletion ownership.
> h2. Steps to Reproduce
> h3. 1. Prepare Hive Metastore
> Start Hive Metastore with a MySQL metadata database and S3A-compatible object
> storage. Confirm that the target database path does not exist.
> For example:
> {code:java}
> s3a://warehouse/hms_direct_race.db
> {code}
> h3. 2. Concurrently invoke create_database
> Use multiple independent Hive Metastore clients and concurrently call:
> {code:java}
> client.createDatabase(database);
> {code}
> All requests must use the same catalog, database name, and location. The
> standalone reproduction used:
> {code:java}
> catalog: hive
> database: hms_direct_race
> concurrency: 16
> {code}
> h3. 3. Observe the Metastore result
> The reproduced result was:
> * 1 request succeeded.
> * 15 requests failed.
> * The failed requests reported a MySQL unique-key conflict:
> {code:java}
> Duplicate entry 'hms_direct_race-hive'
> for key 'DBS.UNIQUE_DATABASE'
> {code}
> * Multiple handlers logged that the same database path was created.
> * Failed requests entered the database-directory cleanup path.
> h3. 4. Observe directory cleanup
> If another request writes content below the database directory before
> rollback cleanup runs, the failed request may move or delete that content
> together with the entire database directory.
> For example, affected Iceberg content may include:
> {code:java}
> iceberg/test1.db/test3/metadata/00000-....metadata.json
> {code}
> When Hadoop Trash is enabled, the path may be moved under a location similar
> to:
> {code:java}
> s3a://warehouse/user/root/.Trash/Current/iceberg/test1.db
> {code}
> h2. Reproduction Log Excerpts
> The following logs are from a standalone reproduction that bypassed all
> higher-level application code. It used 16 independent Hive Metastore clients.
> The run started at {{2026-08-27 02:23:40 UTC}} and targeted database
> {{{}hms_direct_race_20260827T022340Z{}}}.
> h3. Client result
> {code:java}
> HMS_CREATE_DATABASE_RACE host=hive-metastore port=9083
> database=hms_direct_race_20260827T022340Z workers=16
> location=s3a://warehouse/hms_direct_race_20260827T022340Z.db
> worker=00 result=FAILURE error=MetaException: JDODataStoreException:
> Exception thrown flushing changes to datastore Root cause:
> java.sql.SQLIntegrityConstraintViolationException: Duplicate entry
> 'hms_direct_race_20260827t022340z-hive' for key 'DBS.UNIQUE_DATABASE'
> worker=01 result=FAILURE error=MetaException: JDODataStoreException:
> Exception thrown flushing changes to datastore Root cause:
> java.sql.SQLIntegrityConstraintViolationException: Duplicate entry
> 'hms_direct_race_20260827t022340z-hive' for key 'DBS.UNIQUE_DATABASE'
> ...
> worker=11 result=SUCCESS
> ...
> worker=15 result=FAILURE error=MetaException: JDODataStoreException:
> Exception thrown flushing changes to datastore Root cause:
> java.sql.SQLIntegrityConstraintViolationException: Duplicate entry
> 'hms_direct_race_20260827t022340z-hive' for key 'DBS.UNIQUE_DATABASE'
> HMS_CREATE_DATABASE_RACE_COMPLETE elapsed_ms=10380
> {code}
> The final result was one successful request and 15 failures caused by
> {{{}DBS.UNIQUE_DATABASE{}}}.
> h3. HMS audit and directory creation
> Hive Metastore audit logs recorded multiple concurrent {{create_database}}
> requests for the same database. One example is:
> {code:java}
> 2026-08-27T02:23:48,142 INFO [Metastore-Handler-Pool: Thread-1341]
> HiveMetaStore.audit: ugi=root ip=172.19.0.6 cmd=source:172.19.0.6
> create_database: Database(name:hms_direct_race_20260827T022340Z,
> description:direct concurrent create_database test,
> locationUri:s3a://warehouse/hms_direct_race_20260827T022340Z.db,
> parameters:{}, catalogName:hive)
> {code}
> Multiple handlers then processed the same database path:
> {code:java}
> 2026-08-27T02:23:48,401 INFO [Metastore-Handler-Pool: Thread-1340]
> HMSHandler: Creating database path in external directory
> s3a://warehouse/hms_direct_race_20260827T022340Z.db
> 2026-08-27T02:23:48,403 INFO [Metastore-Handler-Pool: Thread-1334]
> HMSHandler: Creating database path in external directory
> s3a://warehouse/hms_direct_race_20260827T022340Z.db
> 2026-08-27T02:23:48,434 INFO [Metastore-Handler-Pool: Thread-1338]
> HMSHandler: Created database path in external directory
> s3a://warehouse/hms_direct_race_20260827T022340Z.db
> 2026-08-27T02:23:48,438 INFO [Metastore-Handler-Pool: Thread-1334]
> HMSHandler: Created database path in external directory
> s3a://warehouse/hms_direct_race_20260827T022340Z.db
> {code}
> h3. Sentinel and cleanup
> A watcher observed the database directory during the race window and uploaded
> a sentinel object below it:
> {code:java}
> WATCHER_START 2026-08-27T02:23:41Z
> DIRECTORY_VISIBLE 2026-08-27T02:23:48Z
> s3://warehouse/hms_direct_race_20260827T022340Z.db
> 46 bytes ->
> `minio/warehouse/hms_direct_race_20260827T022340Z.db/sentinel/metadata/race-sentinel.txt`
> SENTINEL_UPLOAD_DONE 2026-08-27T02:23:48Z
> WATCHER_END 2026-08-27T02:23:48Z
> {code}
> The unique-key conflict and subsequent cleanup appeared together in the Hive
> Metastore logs:
> {code:java}
> Caused by: java.sql.SQLIntegrityConstraintViolationException: Duplicate entry
> 'hms_direct_race_20260827t022340z-hive' for key 'DBS.UNIQUE_DATABASE'
> at
> org.apache.hadoop.hive.metastore.HMSHandler.create_database_core(HMSHandler.java:1274)
> at
> org.apache.hadoop.hive.metastore.HMSHandler.create_database(HMSHandler.java:1343)
> 2026-08-27T02:23:58,482 WARN [Metastore-Handler-Pool: Thread-1342] FileUtils:
> No such file or directory:
> s3a://warehouse/hms_direct_race_20260827T022340Z.db; Force to delete it.
> 2026-08-27T02:23:58,487 ERROR [Metastore-Handler-Pool: Thread-1341]
> FileUtils: Failed to delete
> s3a://warehouse/hms_direct_race_20260827T022340Z.db
> {code}
> After the test, no objects remained under the original path or the
> corresponding Trash path:
> {code:java}
> === OBJECTS UNDER DIRECT PATH ===
> === OBJECTS UNDER TRASH PATH ===
> === DIRECT PATH STAT ===
> mc: <ERROR> Unable to stat
> `minio/warehouse/hms_direct_race_20260827T022340Z.db/`. Object does not exist.
> {code}
> At the same time, the database row still existed in the MySQL {{DBS}} table:
> {code:java}
> DB_ID NAME CTLG_NAME DB_LOCATION_URI
> 104 hms_direct_race_20260827t022340z hive
> s3a://warehouse/hms_direct_race_20260827T022340Z.db
> {code}
> Trash was disabled in this standalone reproduction, so cleanup directly
> deleted the directory instead of moving it to {{{}.Trash/Current{}}}.
> h2. Actual Result
> A failed {{create_database}} request may recursively clean up the entire
> final database directory rather than cleaning up only request-private
> resources.
> This can leave the system in the following state:
> * Database metadata still exists in Hive Metastore.
> * Table metadata may still exist in Hive Metastore.
> * Iceberg {{metadata_location}} values still reference the original
> object-storage paths.
> * The corresponding Iceberg metadata files have been moved to Trash or
> recursively deleted.
> * Subsequent Iceberg queries, startup recovery, or metadata commits receive
> 404/not-found errors.
> * Iceberg tables become unreadable or unwritable.
> h2. Expected Result
> When multiple requests concurrently create the same database:
> * One request should successfully create the database.
> * Other requests should receive {{AlreadyExistsException}} or an equivalent
> database-already-exists error.
> * Failed requests must not delete, move, or otherwise modify the shared
> final database directory.
> * Failed requests may only recursively clean up resources that are provably
> private to that request, such as a request-scoped staging path.
> h2. Root Cause
> {{create_database}} uses a non-atomic check-then-create sequence:
> {code:java}
> get_database_core(...);
> if (database does not exist) {
> create_database_core(...);
> }
> {code}
> Inside {{{}create_database_core{}}}, filesystem ownership is inferred from
> logic similar to:
> {code:java}
> if (!wh.isDir(dbPath)) {
> if (!wh.mkdirs(dbPath)) {
> throw new MetaException(...);
> }
> madeDir = true;
> }
> {code}
> A successful {{mkdirs}} call only shows that the directory became available
> through this request. It does not prove continuing exclusive ownership.
> Multiple requests may set {{madeDir}} to {{{}true{}}}, and even a directory
> created exclusively by one request may subsequently be adopted by a different
> request that wins the metadata transaction.
> When one request commits the Metastore metadata and other requests fail
> because of {{{}DBS.UNIQUE_DATABASE{}}}, the failed requests execute cleanup
> similar to:
> {code:java}
> if (!success) {
> ms.rollbackTransaction();
> if (madeDir) {
> wh.deleteDir(dbPath, true);
> }
> }
> {code}
> The corresponding external database path cleanup in Hive {{4.0.0-beta-1}} is
> equivalent to:
> {code:java}
> if (madeExternalDir && dbExtPath != null) {
> wh.deleteDir(dbExtPath, true, db);
> }
> {code}
> The cleanup scope is the entire final database directory, not an empty
> request-private directory or an object written only by the failed request.
> The unique-key conflict itself is an expected concurrency outcome. The
> data-loss risk comes from treating local directory-creation provenance as
> authorization to recursively delete a shared final location after metadata
> rollback.
> h2. Source References
> * Hive {{4.0.0-beta-1}} {{{}HMSHandler.create_database_core{}}}:
> [source|https://github.com/apache/hive/blob/rel/release-4.0.0-beta-1/standalone-metastore/metastore-server/src/main/java/org/apache/hadoop/hive/metastore/HMSHandler.java#L1138-L1341]
> * Hive {{3.1.2}} {{{}create_database_core{}}}:
> [source|https://github.com/apache/hive/blob/rel/release-3.1.2/standalone-metastore/src/main/java/org/apache/hadoop/hive/metastore/HiveMetaStore.java#L1240-L1287]
> * Hive {{4.2.1}} {{{}HMSHandler.create_database_core{}}}:
> [source|https://github.com/apache/hive/blob/rel/release-4.2.1/standalone-metastore/metastore-server/src/main/java/org/apache/hadoop/hive/metastore/HMSHandler.java#L1163-L1347]
--
This message was sent by Atlassian Jira
(v8.20.10#820010)