wangmingzhou1986 opened a new issue, #68802: URL: https://github.com/apache/doris/issues/68802
### Search before asking - [x] I had searched in the [issues](https://github.com/apache/doris/issues?q=is%3Aissue) and found no similar issues. (Searched `paimon nested_update`, `paimon ClassCastException struct`, `prune nested column paimon jni`. Related but different: #66399, #68214, #68110.) ### Version `doris-4.1.4-rc04-ad35a140c7f` (FE + BE). Paimon external catalog (filesystem catalog on S3/MinIO). Table written by Flink 1.20.5 + Paimon 1.3.1. ### What's Wrong? Reading a Paimon **primary-key, `partial-update`** table whose `ARRAY<ROW<...>>` column uses the **`nested_update`** aggregate function fails through the JNI scanner with a `ClassCastException` when nested column pruning drops the struct field that is the **`nested-key`**. It only fails for splits that need **merge-on-read** (several sorted runs for the key), so it is intermittent: the same query succeeds again once the files are compacted. ``` java.lang.ClassCastException: class org.apache.paimon.data.columnar.heap.HeapBytesVector cannot be cast to class org.apache.paimon.data.columnar.LongColumnVector at org.apache.paimon.data.columnar.VectorizedColumnBatch.getLong(VectorizedColumnBatch.java:85) at org.apache.paimon.data.columnar.ColumnarRow.getLong(ColumnarRow.java:121) at Projection$5236.apply(Unknown Source) at org.apache.paimon.mergetree.compact.aggregate.FieldNestedUpdateAgg.agg(FieldNestedUpdateAgg.java:96) at org.apache.paimon.mergetree.compact.PartialUpdateMergeFunction.updateWithSequenceGroup(PartialUpdateMergeFunction.java:241) at org.apache.paimon.mergetree.compact.PartialUpdateMergeFunction.add(PartialUpdateMergeFunction.java:171) at org.apache.paimon.mergetree.compact.ReducerMergeFunctionWrapper.merge(ReducerMergeFunctionWrapper.java:66) at org.apache.paimon.mergetree.compact.ReducerMergeFunctionWrapper.add(ReducerMergeFunctionWrapper.java:61) at org.apache.paimon.mergetree.compact.SortMergeReaderWithLoserTree$SortMergeIterator.merge(SortMergeReaderWithLoserTree.java:109) at org.apache.paimon.mergetree.compact.SortMergeReaderWithLoserTree$SortMergeIterator.next(SortMergeReaderWithLoserTree.java:97) at org.apache.paimon.mergetree.DropDeleteReader$1.next(DropDeleteReader.java:54) ``` Seen through the MySQL client as `[JNI_ERROR]RuntimeException: java.io.IOException: java.lang.ClassCastException: ...`. Depending on which fields are kept, the message varies, e.g. `ParquetTimestampVector cannot be cast to IntColumnVector`, `HeapBytesVector cannot be cast to IntColumnVector`. **Result matrix** (same table, single primary key, `EXPLAIN` shows `paimonNativeReadSplits=0/1`, i.e. JNI): | fields read from `ROW<ext_id BIGINT, oligo_tag STRING, scheduling_date TIMESTAMP(0)>` (`nested-key` = `ext_id`) | result | |---|---| | `ext_id` | OK | | `ext_id, oligo_tag` | OK | | `ext_id, scheduling_date` | OK | | `oligo_tag` | **ClassCastException** | | `scheduling_date` | **ClassCastException** | | `oligo_tag, scheduling_date` | **ClassCastException** | The same pattern holds for every other `nested_update` column in the table (`ROW<oi_id INT, ...8 fields>`, `ROW<oi_id INT, ..., n1 TIMESTAMP_LTZ, ...>`, `ROW<wo_order_no INT, order_date TIMESTAMP(0)>`): **any projection that keeps the first field (the nested-key) works; any projection without it fails**. With `SET enable_prune_nested_column = false;` (session variable `experimental_enable_prune_nested_column`, default `true`) every query above returns correct results. ### Analysis (please correct me) The stack trace suggests the pruned nested read type is handed to Paimon's merge-on-read path. `FieldNestedUpdateAgg` extracts the nested key with a projection built from the **full** row type (key at position 0), but the rows it receives are already pruned, so position 0 is now `oligo_tag` / `scheduling_date` and `getLong` fails. So either Doris should not prune nested fields that the merge engine needs (the `nested-key` of `nested_update`, and possibly sequence/aggregation inputs) for splits that are not raw-convertible, or Paimon should widen the read type internally. I'm filing here first because Doris decides the read type; happy to cross-post to apache/paimon if you think the fix belongs there. ### What You Expected? Correct values, same as with `enable_prune_nested_column = false`, or the planner keeping the fields required by the merge engine. ### How to Reproduce? 1. In Flink (1.20 + Paimon 1.3.1), create a partial-update table with a `nested_update` array-of-row column: ```sql CREATE TABLE t ( pk INT NOT NULL, exts ARRAY<ROW<ext_id BIGINT, oligo_tag STRING, scheduling_date TIMESTAMP(0)>>, exts_seq BIGINT, PRIMARY KEY (pk) NOT ENFORCED ) WITH ( 'merge-engine' = 'partial-update', 'bucket' = '1', 'fields.exts_seq.sequence-group' = 'exts', 'fields.exts.aggregate-function' = 'nested_update', 'fields.exts.nested-key' = 'ext_id' ); ``` 2. Write the same `pk` in **two or more commits** (e.g. two `INSERT`s with different `ext_id`), and do not let full compaction run, so that the key has several sorted runs. 3. In Doris (Paimon catalog), with default session variables: ```sql SELECT struct_element(e, 'oligo_tag') FROM paimon_catalog.db.t w LATERAL VIEW explode_outer(w.exts) x AS e WHERE w.pk = 1; -- ClassCastException SELECT struct_element(e, 'ext_id'), struct_element(e, 'oligo_tag') FROM paimon_catalog.db.t w LATERAL VIEW explode_outer(w.exts) x AS e WHERE w.pk = 1; -- OK ``` Note: we observed this on a production-like table that is continuously written by a Flink streaming job; on a small, already-compacted table with the same schema the failing query succeeds (consistent with the merge-on-read explanation). The steps above are derived from that observation and the stack trace. ### Anything Else? A second, separate observation (not a bug, mentioned only for context): a query that explodes several such array columns over the whole table fails with `[JNI_ERROR]OutOfMemoryError: Java heap space` under the default `-Xmx1024m` JNI heap; that one is a sizing issue on our side. ### Are you willing to submit PR? - [ ] Yes I am willing to submit a PR! ### Code of Conduct - [x] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
