This is an automated email from the ASF dual-hosted git repository.

yiguolei pushed a commit to branch branch-4.1
in repository https://gitbox.apache.org/repos/asf/doris.git


The following commit(s) were added to refs/heads/branch-4.1 by this push:
     new 872d4e01f7f [fix](regression) Bound large Parquet map output batches 
(#68808)
872d4e01f7f is described below

commit 872d4e01f7f28ee65d7b2b086187fdb452926bb2
Author: Gabriel <[email protected]>
AuthorDate: Fri Oct 9 14:54:24 2026 +0800

    [fix](regression) Bound large Parquet map output batches (#68808)
    
    ### What problem does this PR solve?
    
    Issue Number: N/A
    
    Related PR: N/A
    
    The large-string-map Parquet regression fixture contains two 1 GiB
    string keys in a column chunk larger than 2 GiB, with dictionary and
    PLAIN data pages. Materializing both rows into one output column
    requires a 4 GiB character buffer and can exceed the BE process memory
    limit during ASAN regression runs.
    
    Read this query with `batch_size=1` and disable aggregate pushdown using
    a statement-local hint. Both rows still go through full Map and string
    decoding, while each output batch contains only one large key. The query
    and its expected result remain unchanged. This retains the oversized
    Parquet column-chunk and mixed-encoding coverage; it no longer exercises
    accumulating both keys in one output ColumnString.
    
    ### Release note
    
    None
    
    ### Check List (For Author)
    
    - Test
        - [ ] Regression test
        - [ ] Unit Test
        - [x] Manual test
            - Compiled the modified suite with Groovy 4.0.19 successfully.
            - `git diff --check` passed.
    - Reviewed the batch-size validation and propagation to the native
    Parquet reader on branch-4.1.
    - The HDFS regression was not run locally: the original CI HDFS fixture
    environment is not configured. ASAN end-to-end validation remains for
    CI.
    - Behavior changed:
    - [x] No. Production code is unchanged; the existing test uses
    statement-local execution settings.
        - [ ] Yes.
    - Does this need documentation?
        - [x] No.
        - [ ] Yes.
    
    ### Check List (For Reviewer who merge this PR)
    
    - [ ] Confirm the release note
    - [ ] Confirm test cases
    - [ ] Confirm document
    - [ ] Add branch pick label
---
 .../suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy     | 5 ++++-
 1 file changed, 4 insertions(+), 1 deletion(-)

diff --git 
a/regression-test/suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy 
b/regression-test/suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy
index 7007ecefd3a..d964cb4a225 100644
--- 
a/regression-test/suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy
+++ 
b/regression-test/suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy
@@ -105,7 +105,10 @@ 
suite("test_hdfs_parquet_group0","external,hive,tvf,external_docker") {
 
 
             uri = "${defaultFS}" + 
"/user/doris/tvf_data/test_hdfs_parquet/group0/large_string_map.brotli.parquet"
-            order_qt_test_11 """ select count(arr) from HDFS(
+            // Read both 1 GiB keys one row per batch to avoid a 4 GiB output 
buffer allocation.
+            // Disable aggregate pushdown to retain full decoding of the >2 
GiB column chunk.
+            order_qt_test_11 """ select /*+ SET_VAR(batch_size=1, 
enable_push_down_no_group_agg=false) */
+                        count(arr) from HDFS(
                         "uri" = "${uri}",
                         "hadoop.username" = "${hdfsUserName}",
                         "format" = "parquet"); """


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to