This is an automated email from the ASF dual-hosted git repository.
yiguolei pushed a commit to branch branch-4.1
in repository https://gitbox.apache.org/repos/asf/doris.git
The following commit(s) were added to refs/heads/branch-4.1 by this push:
new 872d4e01f7f [fix](regression) Bound large Parquet map output batches
(#68808)
872d4e01f7f is described below
commit 872d4e01f7f28ee65d7b2b086187fdb452926bb2
Author: Gabriel <[email protected]>
AuthorDate: Fri Oct 9 14:54:24 2026 +0800
[fix](regression) Bound large Parquet map output batches (#68808)
### What problem does this PR solve?
Issue Number: N/A
Related PR: N/A
The large-string-map Parquet regression fixture contains two 1 GiB
string keys in a column chunk larger than 2 GiB, with dictionary and
PLAIN data pages. Materializing both rows into one output column
requires a 4 GiB character buffer and can exceed the BE process memory
limit during ASAN regression runs.
Read this query with `batch_size=1` and disable aggregate pushdown using
a statement-local hint. Both rows still go through full Map and string
decoding, while each output batch contains only one large key. The query
and its expected result remain unchanged. This retains the oversized
Parquet column-chunk and mixed-encoding coverage; it no longer exercises
accumulating both keys in one output ColumnString.
### Release note
None
### Check List (For Author)
- Test
- [ ] Regression test
- [ ] Unit Test
- [x] Manual test
- Compiled the modified suite with Groovy 4.0.19 successfully.
- `git diff --check` passed.
- Reviewed the batch-size validation and propagation to the native
Parquet reader on branch-4.1.
- The HDFS regression was not run locally: the original CI HDFS fixture
environment is not configured. ASAN end-to-end validation remains for
CI.
- Behavior changed:
- [x] No. Production code is unchanged; the existing test uses
statement-local execution settings.
- [ ] Yes.
- Does this need documentation?
- [x] No.
- [ ] Yes.
### Check List (For Reviewer who merge this PR)
- [ ] Confirm the release note
- [ ] Confirm test cases
- [ ] Confirm document
- [ ] Add branch pick label
---
.../suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
diff --git
a/regression-test/suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy
b/regression-test/suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy
index 7007ecefd3a..d964cb4a225 100644
---
a/regression-test/suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy
+++
b/regression-test/suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy
@@ -105,7 +105,10 @@
suite("test_hdfs_parquet_group0","external,hive,tvf,external_docker") {
uri = "${defaultFS}" +
"/user/doris/tvf_data/test_hdfs_parquet/group0/large_string_map.brotli.parquet"
- order_qt_test_11 """ select count(arr) from HDFS(
+ // Read both 1 GiB keys one row per batch to avoid a 4 GiB output
buffer allocation.
+ // Disable aggregate pushdown to retain full decoding of the >2
GiB column chunk.
+ order_qt_test_11 """ select /*+ SET_VAR(batch_size=1,
enable_push_down_no_group_agg=false) */
+ count(arr) from HDFS(
"uri" = "${uri}",
"hadoop.username" = "${hdfsUserName}",
"format" = "parquet"); """
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]