amogh-jahagirdar commented on code in PR #15006:
URL: https://github.com/apache/iceberg/pull/15006#discussion_r2695645099
##########
core/src/main/java/org/apache/iceberg/MergingSnapshotProducer.java:
##########
@@ -1073,6 +1088,139 @@ private List<ManifestFile> newDeleteFilesAsManifests() {
return cachedNewDeleteManifests;
}
+ // Merge duplicates, internally takes care of updating newDeleteFilesBySpec
to remove
+ // duplicates and add the newly merged DV
+ private void mergeDVsAndWrite() {
+ Map<Integer, List<MergedDVContent>> mergedIndicesBySpec =
Maps.newConcurrentMap();
+
+ Tasks.foreach(dvsByReferencedFile.entrySet())
+ .executeWith(ThreadPools.getDeleteWorkerPool())
+ .stopOnFailure()
+ .throwFailureWhenFinished()
+ .run(
+ entry -> {
+ String referencedLocation = entry.getKey();
+ DeleteFileSet dvsToMerge = entry.getValue();
+ // Nothing to merge
+ if (dvsToMerge.size() < 2) {
+ return;
+ }
+
+ MergedDVContent merged = mergePositions(referencedLocation,
dvsToMerge);
+
+ mergedIndicesBySpec
+ .computeIfAbsent(
+ merged.specId, spec ->
Collections.synchronizedList(Lists.newArrayList()))
+ .add(merged);
+ });
+
+ // Update newDeleteFilesBySpec to remove all the duplicates
+ mergedIndicesBySpec.forEach(
+ (specId, mergedDVContent) -> {
+ mergedDVContent.stream()
+ .map(content -> content.mergedDVs)
+ .forEach(duplicateDVs ->
newDeleteFilesBySpec.get(specId).removeAll(duplicateDVs));
+ });
+
+ writeMergedDVs(mergedIndicesBySpec);
+ }
+
+ // Produces a Puffin per partition spec containing the merged DVs for that
spec
Review Comment:
Discussed offline, since it's generally going to be 1 partition spec, we're
OK here, and we can keep the existing behavior. I also realized my dummy spec
idea won't work because the location provider is going to try and use it when
encoding the path with a given tuple so there'd be a misalignment there.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]