+1 for using total rather than average.

On Tue, Sep 29, 2026 at 8:30 AM Steven Wu <[email protected]> wrote:

> Dan,
>
> Here is the slack thread:
> https://apache-iceberg.slack.com/archives/C0BDHBAGARG/p1790113027624689
>
> Yes, your assumption is correct.
>
> Thanks,
> Steven
>
>
>
>
>
> On Tue, Sep 29, 2026 at 8:28 AM Daniel Weeks <[email protected]> wrote:
>
>> Eduard,
>>
>> Could you link to the slack discussion (I wasn't able to find it)?
>>
>> Is it safe to assume that we can determine the average via the
>> combination of *total_bytes* and *value_count*? If that's true, it seems
>> that *total_bytes* would be more valuable for estimation purposes since
>> you have an explicit upper bound on size.
>>
>> -Dan
>>
>> On Tue, Sep 29, 2026 at 8:03 AM Eduard Tudenhöfner <
>> [email protected]> wrote:
>>
>>> Hey everyone,
>>>
>>> We had a few discussions around the *avg_value_size_in_bytes* field on
>>> the Iceberg slack and how it makes e.g. aggregations more difficult than
>>> necessary. We concluded that it's probably best to track the *total*
>>> instead of the *avg.*
>>> That being said, the field is being renamed to *total_bytes* in
>>> https://github.com/apache/iceberg/pull/18308.
>>>
>>> Please speak up if you have any concerns about this change.
>>>
>>> Thanks,
>>> Eduard
>>>
>>

Reply via email to