Ryan's suggestion seems reasonable. Only `totalBytes` metric is stored on
disk, but the Java API also expose `avgValueSize` as a derived value.

But we probably don't need to keep the derived `avgValueSize` value in the
JSON parser (ContentFileParser*)* and REST spec. Since we don't yet support
v4 tables in the REST spec, a breaking change for JSON serialization should
be acceptable.

On Thu, Oct 1, 2026 at 10:23 AM Ryan Blue <[email protected]> wrote:

> Why don't we add `totalBytes` and have `avgValueSize` calculate the value
> from it?
>
> On Thu, Oct 1, 2026 at 7:38 AM Eduard Tudenhöfner <
> [email protected]> wrote:
>
>> It appears the rename in the Java implementation caught bad timing
>> because *avgValueSizes()* was just recently added to a few central
>> places like *ContentFile/Metrics/BaseFile/ContentFileParser *and the
>> REST spec and then released with 1.12.0.
>> Typically we would deprecate those and then introduce *totalBytes()* in
>> all places where *avgValueSizes()* exists today to maintain API
>> stability.
>> Steven brought up the option to replace all *avgValueSizes()* places
>> with the new method and add those API breakages to RevAPI since those
>> methods shouldn't be used by anyone today.
>>
>> I'm raising this here to get people's thoughts on this
>>
>>
>>
>> On Wed, Sep 30, 2026 at 12:40 AM Ryan Blue <[email protected]> wrote:
>>
>>> +1 for using total rather than average.
>>>
>>> On Tue, Sep 29, 2026 at 8:30 AM Steven Wu <[email protected]> wrote:
>>>
>>>> Dan,
>>>>
>>>> Here is the slack thread:
>>>> https://apache-iceberg.slack.com/archives/C0BDHBAGARG/p1790113027624689
>>>>
>>>> Yes, your assumption is correct.
>>>>
>>>> Thanks,
>>>> Steven
>>>>
>>>>
>>>>
>>>>
>>>>
>>>> On Tue, Sep 29, 2026 at 8:28 AM Daniel Weeks <[email protected]> wrote:
>>>>
>>>>> Eduard,
>>>>>
>>>>> Could you link to the slack discussion (I wasn't able to find it)?
>>>>>
>>>>> Is it safe to assume that we can determine the average via the
>>>>> combination of *total_bytes* and *value_count*? If that's true, it
>>>>> seems that *total_bytes* would be more valuable for estimation
>>>>> purposes since you have an explicit upper bound on size.
>>>>>
>>>>> -Dan
>>>>>
>>>>> On Tue, Sep 29, 2026 at 8:03 AM Eduard Tudenhöfner <
>>>>> [email protected]> wrote:
>>>>>
>>>>>> Hey everyone,
>>>>>>
>>>>>> We had a few discussions around the *avg_value_size_in_bytes* field
>>>>>> on the Iceberg slack and how it makes e.g. aggregations more difficult 
>>>>>> than
>>>>>> necessary. We concluded that it's probably best to track the *total*
>>>>>> instead of the *avg.*
>>>>>> That being said, the field is being renamed to *total_bytes* in
>>>>>> https://github.com/apache/iceberg/pull/18308.
>>>>>>
>>>>>> Please speak up if you have any concerns about this change.
>>>>>>
>>>>>> Thanks,
>>>>>> Eduard
>>>>>>
>>>>>

Reply via email to