jayzhan211 commented on code in PR #25877:
URL: https://github.com/apache/datafusion/pull/25877#discussion_r4156716698


##########
datafusion/functions-aggregate/src/count.rs:
##########
@@ -824,6 +838,125 @@ impl GroupsAccumulator for CountGroupsAccumulator {
     }
 }
 
+/// [`CountGroupsAccumulator`] with the counts stored in blocks of
+/// `block_size` groups, so growing past a block never reallocates and
+/// copies the existing counts.
+#[derive(Debug)]
+struct BlockedCountGroupsAccumulator {
+    /// Count per group, see [`CountGroupsAccumulator::counts`].
+    counts: BlockedVec<i64>,
+}
+
+impl BlockedCountGroupsAccumulator {
+    fn new(block_size: usize) -> Self {
+        Self {
+            counts: BlockedVec::new(block_size),
+        }
+    }
+
+    fn counts_to_array(counts: Vec<i64>) -> ArrayRef {
+        // zero copy, count is never null
+        Arc::new(Int64Array::new(counts.into(), None))
+    }
+
+    fn blocks_to_arrays(blocks: impl IntoIterator<Item = Vec<i64>>) -> 
Vec<ArrayRef> {
+        blocks.into_iter().map(Self::counts_to_array).collect()
+    }
+
+    fn take(&mut self, emit_to: BlockedEmitTo) -> Vec<ArrayRef> {
+        match emit_to {
+            BlockedEmitTo::All => 
Self::blocks_to_arrays(self.counts.take_all()),
+            BlockedEmitTo::NextBlock => {
+                Self::blocks_to_arrays(self.counts.take_next_block())
+            }
+            BlockedEmitTo::First(n) => {
+                Self::blocks_to_arrays([self.counts.take_first(n)])
+            }
+        }
+    }
+}
+
+impl BlockedGroupsAccumulator for BlockedCountGroupsAccumulator {
+    fn block_size(&self) -> usize {
+        self.counts.block_size()
+    }
+
+    fn update_batch(
+        &mut self,
+        values: &[ArrayRef],
+        group_indices: &[BlocksIndex],
+        opt_filter: Option<&BooleanArray>,
+        total_num_groups: usize,
+    ) -> Result<()> {
+        assert_eq!(values.len(), 1, "single argument to update_batch");
+        let values = &values[0];
+        let nulls = values.logical_nulls().filter(|n| n.null_count() > 0);
+
+        self.counts.grow_to(total_num_groups, 0);
+
+        // Add one to each group's counter for each non null, non filtered 
value
+        // SAFETY: group_index is guaranteed to be in bounds and less than 
total_num_groups
+        unsafe {
+            self.counts.update_unchecked(

Review Comment:
   `update_batch` calls `update_unchecked` with only a `SAFETY` comment about 
the caller's contract. That means a safe public path 
(`create_blocked_groups_accumulator` → `update_batch`) can write out of bounds. 
Passing `BlocksIndex::new(0, 1)` with `total_num_groups = 1` aborts in debug 
with `unsafe precondition(s) violated: slice::get_unchecked_mut` at 
`blocked_vec.rs:282`. `merge_batch` already goes through the checked 
`update_with`, and `all_in_bounds` is one vectorized pass, so the checked path 
should cost little here too. The flat count on main has the same pattern, so 
this is fine to handle in a follow-up if you'd rather.
   
   ```diff
   -        self.counts.grow_to(total_num_groups, 0);
   -
   -        // Add one to each group's counter for each non null, non filtered 
value
   -        // SAFETY: group_index is guaranteed to be in bounds and less than 
total_num_groups
   -        unsafe {
   -            self.counts.update_unchecked(
   -                group_indices,
   -                nulls.as_ref(),
   -                opt_filter,
   -                |count| *count += 1,
   -            );
   -        }
   +        // Add one to each group's counter for each non null, non filtered 
value
   +        self.counts.update(
   +            total_num_groups,
   +            0,
   +            group_indices,
   +            nulls.as_ref(),
   +            opt_filter,
   +            |count| *count += 1,
   +        );
   ```



##########
datafusion/physical-plan/src/aggregates/hash_stream.rs:
##########
@@ -316,20 +316,44 @@ impl PartialHashAggregateStream {
                         break;
                     }
                     HandleInputResult::OOM => {
-                        let materialized_group_states = 
hash_table.take_state_batch()?.ok_or_else(|| {
-                            internal_datafusion_err!(
+                        let materialized_group_states =
+                            hash_table.take_state_batches()?;
+                        if materialized_group_states.is_empty() {
+                            return Err(internal_datafusion_err!(
                                 "Partial hash aggregate ran out of memory with 
no aggregated groups"
-                            )
-                        })?;
+                            ));
+                        }
 
                         self.early_emit_count.add(1);
                         timer.done();
-                        self.emit_on_memory_pressure(
-                            materialized_group_states,
-                            &mut emitter,
-                            hash_table.memory_size(),
-                        )
-                        .await?;
+
+                        // Blocked storage returns one batch per block, moved 
out of
+                        // the table without copying. Emit them in turn, 
keeping the
+                        // batches not emitted yet in the reservation.
+                        let pending_memory: usize = materialized_group_states
+                            .iter()
+                            // Don't include the first batch since we will 
emit it right away and release its memory
+                            // if it needs slicing then a we will try to hold 
on that reservation while slicing
+                            .skip(1)
+                            .map(|b| b.memory_size)
+                            .sum();
+
+                        // Make sure we can hold on the hash tables and all 
the batches that need to be emitted (except the first one)
+                        // if we can't hold it than we can't do anything about 
it.
+                        self.reservation.try_resize(

Review Comment:
   `try_resize(hash_table.memory_size() + pending_memory)?` dropped the 
fallback main had: when holding the materialized state doesn't fit, main 
reserved only the table and emitted unreserved. With blocked storage (a blocked 
`count` is enough, even with flat keys), partial aggregation now fails with 
ResourcesExhausted. 4 existing tests fail here and pass on the merge-base 
6924c6a: both `aggregate_grouping_sets_*_with_spill` (`Failed to allocate 
additional 872.0 B ... pool_size: 500.0 B`), 
`test_partial_hash_stream_emits_whole_batch_when_held_batch_does_not_fit`, and 
`test_partial_hash_stream_accounts_held_batch_on_memory_pressure_while_slicing`.
   
   Fix (with it, both grouping-sets tests pass locally):
   ```diff
   -                        self.reservation.try_resize(
   -                            hash_table.memory_size()
   -                              // The batches that need to be held while 
emitting each one
   -                              + pending_memory,
   -                        )?;
   +                        let hold_pending = match self
   +                            .reservation
   +                            .try_resize(hash_table.memory_size() + 
pending_memory)
   +                        {
   +                            Ok(()) => true,
   +                            // Can't hold the pending blocks: emit them 
unreserved, like
   +                            // the flat path emits the whole batch
   +                            Err(DataFusionError::ResourcesExhausted(_)) => {
   +                                
self.reservation.try_resize(hash_table.memory_size())?;
   +                                false
   +                            }
   +                            Err(e) => return Err(e),
   +                        };
   
                            for (i, batch) in
                                
materialized_group_states.into_iter().enumerate()
                            {
   -                            if i != 0 {
   +                            if i != 0 && hold_pending {
                                    
self.reservation.try_shrink(batch.memory_size)?;
                                }
   ```
   
   The two `hash_stream` tests also assume a single state batch 
(`first.num_rows() == num_groups`, and a shared `data_ptr` between the first 
two outputs). Please update them for one batch per block, e.g. assert that the 
total rows match and that `reserved` covers the blocks not emitted yet.



##########
datafusion/physical-plan/src/aggregates/group_values/blocked/primitive.rs:
##########
@@ -0,0 +1,412 @@
+// Licensed to the Apache Software Foundation (ASF) under one
+// or more contributor license agreements.  See the NOTICE file
+// distributed with this work for additional information
+// regarding copyright ownership.  The ASF licenses this file
+// to you under the Apache License, Version 2.0 (the
+// "License"); you may not use this file except in compliance
+// with the License.  You may obtain a copy of the License at
+//
+//   http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing,
+// software distributed under the License is distributed on an
+// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+// KIND, either express or implied.  See the License for the
+// specific language governing permissions and limitations
+// under the License.
+
+use std::mem::size_of;
+use std::sync::Arc;
+
+use arrow::array::{
+    ArrayRef, ArrowNativeTypeOp, ArrowPrimitiveType, NullBufferBuilder, 
PrimitiveArray,
+    cast::AsArray,
+};
+use arrow::datatypes::DataType;
+use datafusion_common::Result;
+use datafusion_common::hash_utils::RandomState;
+use datafusion_expr::{BlockedEmitTo, BlocksIndex};
+use 
datafusion_functions_aggregate_common::aggregate::groups_accumulator::blocked_vec::BlockedVec;
+use hashbrown::hash_table::HashTable;
+
+use super::BlockedGroupValues;
+use crate::aggregates::group_values::HashValue;
+
+/// A [`BlockedGroupValues`] storing a single column of primitive values.
+///
+/// Like `GroupValuesPrimitive`, but the values are stored in a [`BlockedVec`]
+/// and the hash table entries hold `(group_index, value)` instead of
+/// `(group_index, hash)`, so probing never reads the (blocked) values.
+pub(crate) struct BlockedGroupValuesPrimitive<T: ArrowPrimitiveType> {
+    /// The data type of the output array
+    data_type: DataType,
+    /// Stores `(group_index, value)` based on the hash of the value
+    ///
+    /// Storing the value instead of its hash means comparing keys never goes
+    /// through `values`, and rehashing is cheap for primitive values.
+    map: HashTable<(BlocksIndex, T::Native)>,
+    /// The group index of the null value if any
+    null_group: Option<BlocksIndex>,
+    /// The values for each group index
+    values: BlockedVec<T::Native>,
+    /// The random state used to generate hashes
+    random_state: RandomState,
+}
+
+impl<T: ArrowPrimitiveType> BlockedGroupValuesPrimitive<T> {
+    pub(crate) fn new(data_type: DataType, block_size: usize) -> Self {
+        assert!(PrimitiveArray::<T>::is_compatible(&data_type));
+        Self {
+            data_type,
+            map: HashTable::with_capacity(128),
+            values: BlockedVec::new(block_size),
+            null_group: None,
+            random_state: crate::aggregates::AGGREGATION_HASH_SEED,
+        }
+    }
+
+    /// Builds the output array of one block whose first group is `start`
+    fn build_block(&self, values: Vec<T::Native>, start: usize) -> ArrayRef {
+        let len = values.len();
+        let block_size = self.values.block_size();
+        let nulls = self
+            .null_group
+            .map(|null_group| null_group.flat(block_size))
+            .filter(|null_group| (start..start + len).contains(null_group))
+            .map(|null_group| {
+                let null_idx = null_group - start;
+                let mut buffer = NullBufferBuilder::new(len);
+                buffer.append_n_non_nulls(null_idx);
+                buffer.append_null();
+                buffer.append_n_non_nulls(len - null_idx - 1);
+                // NOTE: The inner builder must be constructed as there is at 
least one null
+                buffer.finish().unwrap()
+            });
+        Arc::new(
+            PrimitiveArray::<T>::new(values.into(), nulls)
+                .with_data_type(self.data_type.clone()),
+        )
+    }
+
+    /// Updates the null group after the first `n` groups were emitted
+    fn shift_null_group(&mut self, n: usize) {
+        let block_size = self.values.block_size();
+        self.null_group = self
+            .null_group
+            .and_then(|null_group| null_group.flat(block_size).checked_sub(n))
+            .map(|flat| BlocksIndex::from_flat(flat, block_size));
+    }
+}
+
+impl<T: ArrowPrimitiveType> BlockedGroupValues for 
BlockedGroupValuesPrimitive<T>
+where
+    T::Native: HashValue,
+{
+    fn intern(&mut self, cols: &[ArrayRef], groups: &mut Vec<BlocksIndex>) -> 
Result<()> {
+        assert_eq!(cols.len(), 1);
+        groups.clear();
+
+        for v in cols[0].as_primitive::<T>() {
+            let group_id = match v {
+                None => *self
+                    .null_group
+                    .get_or_insert_with(|| 
self.values.push(Default::default())),
+                Some(key) => {
+                    // Fold equivalence-class duplicates (e.g. `-0.0` → `+0.0`)
+                    // so the bit-equal `is_eq` matches and the stored value is
+                    // the canonical representative.
+                    let key = key.canonicalize();
+                    let state = &self.random_state;
+                    let hash = key.hash(state);
+                    let insert = self.map.entry(
+                        hash,
+                        |&(_, k)| k.is_eq(key),
+                        |&(_, k)| k.hash(state),
+                    );
+
+                    match insert {
+                        hashbrown::hash_table::Entry::Occupied(o) => o.get().0,
+                        hashbrown::hash_table::Entry::Vacant(v) => {
+                            let g = self.values.push(key);
+                            v.insert((g, key));
+                            g
+                        }
+                    }
+                }
+            };
+            groups.push(group_id)
+        }
+        Ok(())
+    }
+
+    fn size(&self) -> usize {
+        self.map.capacity() * size_of::<(BlocksIndex, T::Native)>()
+            + self.values.allocated_size()
+    }
+
+    fn is_empty(&self) -> bool {
+        self.values.is_empty()
+    }
+
+    fn len(&self) -> usize {
+        self.values.len()
+    }
+
+    fn emit(&mut self, emit_to: BlockedEmitTo) -> Result<Vec<Vec<ArrayRef>>> {
+        let blocks = match emit_to {
+            BlockedEmitTo::All => {
+                self.map.clear();
+                let mut start = 0;
+                let blocks = self
+                    .values
+                    .take_all()
+                    .into_iter()
+                    .map(|block| {
+                        let len = block.len();
+                        let array = self.build_block(block, start);
+                        start += len;
+                        vec![array]
+                    })
+                    .collect();
+                self.null_group = None;
+                blocks
+            }
+            BlockedEmitTo::NextBlock => {

Review Comment:
   `NextBlock` runs `map.retain` over every entry to renumber it, so draining n 
groups one block at a time costs O(n²/block_size). With 10M groups and 8192-row 
blocks that's about 1.2k passes over the full map. Incremental emit is what 
`NextBlock` is for, so this will show up once ordered or external callers use 
it. Nothing calls it yet, so a follow-up is fine.
   
   A suggestion: store absolute block numbers in the map plus a `first_block` 
offset. `intern` subtracts the offset when it returns an index and adds it when 
it inserts one. `NextBlock` then removes only the emitted block's keys by 
lookup, at O(block_size) per block:
   ```rs
   BlockedEmitTo::NextBlock => {
       let Some(block) = self.values.take_next_block() else {
           return Ok(vec![]);
       };
       let first_block = self.first_block;
       let state = &self.random_state;
       for &key in &block {
           // The null group's slot holds `Default`, so match on the block too
           if let Ok(entry) = self.map.find_entry(key.hash(state), |&(g, k)| {
               g.block_index() == first_block && k.is_eq(key)
           }) {
               entry.remove();
           }
       }
       self.first_block += 1;
       let len = block.len();
       let array = self.build_block(block, 0);
       self.shift_null_group(len);
       vec![vec![array]]
   }
   ```
   `All` and `clear_shrink` reset `first_block` to 0. `First(n)` would also 
need to account for the offset.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to