Doris-Breakwater commented on issue #68802:
URL: https://github.com/apache/doris/issues/68802#issuecomment-6072421979

   Breakwater-GitHub-Analysis-Slot: slot_07af1467818c
   
   **Initial assessment:** The reported projection/merge-on-read explanation is 
strongly supported by the source at the exact Doris commit 
`ad35a140c7fd0b842f18c23300bac581f7d04326`. This should be investigated as a 
Paimon JNI nested-projection correctness issue. It is currently open and 
unlabelled; Paimon/external-catalog maintainers are the appropriate first 
owners. I have verified the code mechanism, but have not run the Flink → Paimon 
→ Doris reproducer, so the deployed failure and fix remain to be confirmed 
experimentally.
   
   **Verified source evidence**
   
   - Doris 
[`PaimonReadTypeProjection`](https://github.com/apache/doris/blob/ad35a140c7fd0b842f18c23300bac581f7d04326/fe/be-java-extensions/paimon-connector/src/main/java/org/apache/doris/paimon/PaimonReadTypeProjection.java#L57)
 rebuilds ROW fields from the requested children and recursively applies that 
shape through ARRAY. It preserves field IDs, but does not inspect aggregation 
options or retain a `nested-key` automatically. 
[`PaimonJniScanner.initReader()`](https://github.com/apache/doris/blob/ad35a140c7fd0b842f18c23300bac581f7d04326/fe/be-java-extensions/paimon-connector/src/main/java/org/apache/doris/paimon/PaimonJniScanner.java#L196)
 passes this shape to `withReadType()` before creating the reader.
   - In Paimon 1.4.2, 
[`MergeFileSplitRead.withReadType()`](https://github.com/apache/paimon/blob/release-1.4.2/paimon-core/src/main/java/org/apache/paimon/operation/MergeFileSplitRead.java#L133)
 applies the merge factory's read-type adjustment and uses the resulting type 
for file reading and merging. 
[`PartialUpdateMergeFunction.Factory`](https://github.com/apache/paimon/blob/release-1.4.2/paimon-core/src/main/java/org/apache/paimon/mergetree/compact/PartialUpdateMergeFunction.java#L492)
 remaps aggregators at the top level while retaining suppliers created from the 
original field types. Its adjustment adds missing sequence-group fields, but 
does not restore nested aggregation inputs.
   - 
[`FieldNestedUpdateAgg`](https://github.com/apache/paimon/blob/release-1.4.2/paimon-core/src/main/java/org/apache/paimon/mergetree/compact/aggregate/FieldNestedUpdateAgg.java#L57)
 builds its key projection from that original element ROW type. During 
aggregation it applies the projection to the incoming rows. For your schema, 
dropping `ext_id` makes `oligo_tag` occupy position 0, while the generated key 
accessor still expects BIGINT there. Paimon's columnar `getLong()` casts the 
vector at that position to `LongColumnVector`, explaining the reported 
exception.
   
   This also explains why the first input can succeed and the error appears 
when arrays actually need aggregation. The reported compaction dependence is 
consistent with this mechanism, although JNI selection alone does not prove 
that a split requires merging. Keeping the key at its original position 
explains your successful matrix entries; it is not a general guarantee for 
other schemas.
   
   The relevant Doris integration change is 
[#66573](https://github.com/apache/doris/pull/66573). Its nested-pruning 
regression fixture is append-only, so it does not exercise this merge 
dependency. The references listed in the issue do not establish a fix for this 
path.
   
   **Version detail and evidence still needed**
   
   The exact Doris commit declares **Paimon 1.4.2** in 
[`fe/pom.xml`](https://github.com/apache/doris/blob/ad35a140c7fd0b842f18c23300bac581f7d04326/fe/pom.xml#L391).
 The Flink writer's **1.3.1** version does not identify the JNI reader version. 
Please provide:
   
   1. Actual deployed Paimon reader jar versions, any replacement jars, and 
confirmation that all participating FE/BE nodes use the reported build.
   2. A self-contained reproduction with the exact INSERT statements, sequence 
values, commit/checkpoint boundaries, and expected rows. Preserve a snapshot 
with overlapping runs; two commits alone do not guarantee merge work survives 
compaction.
   3. Full BE Java stack trace and query ID, plus EXPLAIN output showing pruned 
types/access paths for a failing query and the same query with pruning 
disabled. Include the table's effective options, schema ID, snapshot ID, 
physical file format, and split/file overlap metadata. Redact credentials and 
private storage locations.
   
   **Workaround and next steps**
   
   Use the already validated session workaround:
   
   ```sql
   SET enable_prune_nested_column = false;
   ```
   
   At this commit the variable is experimental and defaults to true. Adding the 
nested key to SELECT is a schema-specific workaround; compaction does not 
provide a durable remedy under continuing writes.
   
   For maintainers, first reproduce against the deployed reader version on a 
fixed snapshot. A conservative Doris mitigation is to prevent nested pruning of 
affected merge-dependent columns while retaining top-level projection. A Paimon 
fix could retain the required nested schema before merging and project it 
afterwards. Both require preserving field order, types, and the JNI output 
schema: simply widening the scanner read type is insufficient because Doris 
decodes using its requested shape. Paimon's existing outer projection is 
top-level only, so nested output projection also needs explicit handling.
   
   Regression coverage should compare actual values with pruning disabled, 
using overlapping runs and compacted controls, omitted/included keys, 
non-first/composite keys, and same-typed fields that could hide an ordinal 
mismatch without throwing. Verify sequence-group ordering and duplicate 
nested-key updates too. The separate Java heap exception should be investigated 
independently; it does not explain this typed-accessor failure.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to