Doris-Breakwater commented on issue #68802: URL: https://github.com/apache/doris/issues/68802#issuecomment-6072421979
Breakwater-GitHub-Analysis-Slot: slot_07af1467818c **Initial assessment:** The reported projection/merge-on-read explanation is strongly supported by the source at the exact Doris commit `ad35a140c7fd0b842f18c23300bac581f7d04326`. This should be investigated as a Paimon JNI nested-projection correctness issue. It is currently open and unlabelled; Paimon/external-catalog maintainers are the appropriate first owners. I have verified the code mechanism, but have not run the Flink → Paimon → Doris reproducer, so the deployed failure and fix remain to be confirmed experimentally. **Verified source evidence** - Doris [`PaimonReadTypeProjection`](https://github.com/apache/doris/blob/ad35a140c7fd0b842f18c23300bac581f7d04326/fe/be-java-extensions/paimon-connector/src/main/java/org/apache/doris/paimon/PaimonReadTypeProjection.java#L57) rebuilds ROW fields from the requested children and recursively applies that shape through ARRAY. It preserves field IDs, but does not inspect aggregation options or retain a `nested-key` automatically. [`PaimonJniScanner.initReader()`](https://github.com/apache/doris/blob/ad35a140c7fd0b842f18c23300bac581f7d04326/fe/be-java-extensions/paimon-connector/src/main/java/org/apache/doris/paimon/PaimonJniScanner.java#L196) passes this shape to `withReadType()` before creating the reader. - In Paimon 1.4.2, [`MergeFileSplitRead.withReadType()`](https://github.com/apache/paimon/blob/release-1.4.2/paimon-core/src/main/java/org/apache/paimon/operation/MergeFileSplitRead.java#L133) applies the merge factory's read-type adjustment and uses the resulting type for file reading and merging. [`PartialUpdateMergeFunction.Factory`](https://github.com/apache/paimon/blob/release-1.4.2/paimon-core/src/main/java/org/apache/paimon/mergetree/compact/PartialUpdateMergeFunction.java#L492) remaps aggregators at the top level while retaining suppliers created from the original field types. Its adjustment adds missing sequence-group fields, but does not restore nested aggregation inputs. - [`FieldNestedUpdateAgg`](https://github.com/apache/paimon/blob/release-1.4.2/paimon-core/src/main/java/org/apache/paimon/mergetree/compact/aggregate/FieldNestedUpdateAgg.java#L57) builds its key projection from that original element ROW type. During aggregation it applies the projection to the incoming rows. For your schema, dropping `ext_id` makes `oligo_tag` occupy position 0, while the generated key accessor still expects BIGINT there. Paimon's columnar `getLong()` casts the vector at that position to `LongColumnVector`, explaining the reported exception. This also explains why the first input can succeed and the error appears when arrays actually need aggregation. The reported compaction dependence is consistent with this mechanism, although JNI selection alone does not prove that a split requires merging. Keeping the key at its original position explains your successful matrix entries; it is not a general guarantee for other schemas. The relevant Doris integration change is [#66573](https://github.com/apache/doris/pull/66573). Its nested-pruning regression fixture is append-only, so it does not exercise this merge dependency. The references listed in the issue do not establish a fix for this path. **Version detail and evidence still needed** The exact Doris commit declares **Paimon 1.4.2** in [`fe/pom.xml`](https://github.com/apache/doris/blob/ad35a140c7fd0b842f18c23300bac581f7d04326/fe/pom.xml#L391). The Flink writer's **1.3.1** version does not identify the JNI reader version. Please provide: 1. Actual deployed Paimon reader jar versions, any replacement jars, and confirmation that all participating FE/BE nodes use the reported build. 2. A self-contained reproduction with the exact INSERT statements, sequence values, commit/checkpoint boundaries, and expected rows. Preserve a snapshot with overlapping runs; two commits alone do not guarantee merge work survives compaction. 3. Full BE Java stack trace and query ID, plus EXPLAIN output showing pruned types/access paths for a failing query and the same query with pruning disabled. Include the table's effective options, schema ID, snapshot ID, physical file format, and split/file overlap metadata. Redact credentials and private storage locations. **Workaround and next steps** Use the already validated session workaround: ```sql SET enable_prune_nested_column = false; ``` At this commit the variable is experimental and defaults to true. Adding the nested key to SELECT is a schema-specific workaround; compaction does not provide a durable remedy under continuing writes. For maintainers, first reproduce against the deployed reader version on a fixed snapshot. A conservative Doris mitigation is to prevent nested pruning of affected merge-dependent columns while retaining top-level projection. A Paimon fix could retain the required nested schema before merging and project it afterwards. Both require preserving field order, types, and the JNI output schema: simply widening the scanner read type is insufficient because Doris decodes using its requested shape. Paimon's existing outer projection is top-level only, so nested output projection also needs explicit handling. Regression coverage should compare actual values with pruning disabled, using overlapping runs and compacted controls, omitted/included keys, non-first/composite keys, and same-typed fields that could hide an ordinal mismatch without throwing. Verify sequence-group ordering and duplicate nested-key updates too. The separate Java heap exception should be investigated independently; it does not explain this typed-accessor failure. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
