github-actions[bot] commented on code in PR #68808:
URL: https://github.com/apache/doris/pull/68808#discussion_r4226410447


##########
regression-test/suites/external_table_p0/tvf/test_hdfs_parquet_group0.groovy:
##########
@@ -105,7 +105,10 @@ 
suite("test_hdfs_parquet_group0","external,hive,tvf,external_docker") {
 
 
             uri = "${defaultFS}" + 
"/user/doris/tvf_data/test_hdfs_parquet/group0/large_string_map.brotli.parquet"
-            order_qt_test_11 """ select count(arr) from HDFS(
+            // Read both 1 GiB keys one row per batch to avoid a 4 GiB output 
buffer allocation.
+            // Disable aggregate pushdown to retain full decoding of the >2 
GiB column chunk.
+            order_qt_test_11 """ select /*+ SET_VAR(batch_size=1, 
enable_push_down_no_group_agg=false) */

Review Comment:
   [P2] Honor this row cap with V2 adaptive batching disabled. With the default 
`enable_file_scanner_v2=true`, setting the mutable BE 
`enable_adaptive_batch_size=false` skips V2's `set_batch_size` calls. 
`TableReader` remains at zero and native Parquet reads its 4096-row default 
despite this `batch_size=1` hint. Both 1 GiB map keys can then enter one output 
batch, recreating the allocation this test is meant to avoid. Seed V2's reader 
cap from the query batch size even when adaptive batching is off, or constrain 
this regression to a scanner path that honors the hint.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to