claudevdm commented on code in PR #40252:
URL: https://github.com/apache/beam/pull/40252#discussion_r4126859491
##########
sdks/java/io/iceberg/src/main/java/org/apache/beam/sdk/io/iceberg/AddFiles.java:
##########
@@ -942,16 +950,16 @@ private String getPartitionFromFilePath(String filePath) {
* <p>In these cases, we output the DataFile to the DLQ, because assigning
an incorrect
* partition may lead to it being incorrectly ignored by downstream
queries.
*/
- static String getPartitionFromMetrics(
+ static PartitionKey getPartitionFromMetrics(
Metrics metrics, InputFile inputFile, Table table, @Nullable
ParquetMetadata preReadFooter)
throws UnknownPartitionException {
List<PartitionField> fields = table.spec().fields();
List<Integer> sourceIds =
fields.stream().map(PartitionField::sourceId).collect(Collectors.toList());
Metrics partitionMetrics;
// Check if metrics already includes partition columns (configured by
table properties):
- if (metrics.lowerBounds().keySet().containsAll(sourceIds)
- && metrics.upperBounds().keySet().containsAll(sourceIds)) {
+ if (orEmpty(metrics.lowerBounds()).keySet().containsAll(sourceIds)
+ && orEmpty(metrics.upperBounds()).keySet().containsAll(sourceIds)) {
partitionMetrics = metrics;
} else {
Review Comment:
The fast path now also requires full metrics mode for every partition source
column. Otherwise we recompute those columns in full mode from the footer,
which is already in memory.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]