rdblue commented on code in PR #5981:
URL: https://github.com/apache/iceberg/pull/5981#discussion_r996153635
##########
core/src/main/java/org/apache/iceberg/ReachableFileCleanup.java:
##########
@@ -85,19 +79,60 @@ public void cleanFiles(TableMetadata beforeExpiration,
TableMetadata afterExpira
}
private Set<ManifestFile> readManifests(Set<Snapshot> snapshots) {
- Set<ManifestFile> manifestFiles = Sets.newHashSet();
- for (Snapshot snapshot : snapshots) {
- try (CloseableIterable<ManifestFile> manifestFilesForSnapshot =
readManifestFiles(snapshot)) {
- for (ManifestFile manifestFile : manifestFilesForSnapshot) {
- manifestFiles.add(manifestFile.copy());
- }
- } catch (IOException e) {
- throw new RuntimeIOException(
- e, "Failed to close manifest list: %s",
snapshot.manifestListLocation());
- }
- }
+ Set<ManifestFile> manifests = ConcurrentHashMap.newKeySet();
+ Tasks.foreach(snapshots)
+ .retry(3)
+ .stopOnFailure()
+ .throwFailureWhenFinished()
+ .executeWith(planExecutorService)
+ .onFailure(
+ (snapshot, exc) ->
+ LOG.warn(
+ "Failed to determine manifests for snapshot {}",
snapshot.snapshotId(), exc))
+ .run(
+ snapshot -> {
+ try (CloseableIterable<ManifestFile> manifestFilesForSnapshot =
+ readManifestFiles(snapshot)) {
+ for (ManifestFile manifestFile : manifestFilesForSnapshot) {
+ manifests.add(manifestFile.copy());
+ }
+ } catch (IOException e) {
+ throw new RuntimeIOException(
+ e, "Failed to close manifest list: %s",
snapshot.manifestListLocation());
+ }
+ });
+
+ return manifests;
+ }
+
+ private Set<ManifestFile> manifestFilesToDelete(
+ Set<ManifestFile> currentManifests, Set<Snapshot> expiredSnapshots) {
Review Comment:
Why not use the same idea as the file reachability method? We expect the
number of current manifests to be much larger than the number of expired
manifests, so keeping the expired set in memory is a better option. In
addition, if we eliminate all of the expired manifests, we can exit early
rather than reading all of the current manifests.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]