Ryan's suggestion seems reasonable. Only `totalBytes` metric is stored on disk, but the Java API also expose `avgValueSize` as a derived value.
But we probably don't need to keep the derived `avgValueSize` value in the JSON parser (ContentFileParser*)* and REST spec. Since we don't yet support v4 tables in the REST spec, a breaking change for JSON serialization should be acceptable. On Thu, Oct 1, 2026 at 10:23 AM Ryan Blue <[email protected]> wrote: > Why don't we add `totalBytes` and have `avgValueSize` calculate the value > from it? > > On Thu, Oct 1, 2026 at 7:38 AM Eduard Tudenhöfner < > [email protected]> wrote: > >> It appears the rename in the Java implementation caught bad timing >> because *avgValueSizes()* was just recently added to a few central >> places like *ContentFile/Metrics/BaseFile/ContentFileParser *and the >> REST spec and then released with 1.12.0. >> Typically we would deprecate those and then introduce *totalBytes()* in >> all places where *avgValueSizes()* exists today to maintain API >> stability. >> Steven brought up the option to replace all *avgValueSizes()* places >> with the new method and add those API breakages to RevAPI since those >> methods shouldn't be used by anyone today. >> >> I'm raising this here to get people's thoughts on this >> >> >> >> On Wed, Sep 30, 2026 at 12:40 AM Ryan Blue <[email protected]> wrote: >> >>> +1 for using total rather than average. >>> >>> On Tue, Sep 29, 2026 at 8:30 AM Steven Wu <[email protected]> wrote: >>> >>>> Dan, >>>> >>>> Here is the slack thread: >>>> https://apache-iceberg.slack.com/archives/C0BDHBAGARG/p1790113027624689 >>>> >>>> Yes, your assumption is correct. >>>> >>>> Thanks, >>>> Steven >>>> >>>> >>>> >>>> >>>> >>>> On Tue, Sep 29, 2026 at 8:28 AM Daniel Weeks <[email protected]> wrote: >>>> >>>>> Eduard, >>>>> >>>>> Could you link to the slack discussion (I wasn't able to find it)? >>>>> >>>>> Is it safe to assume that we can determine the average via the >>>>> combination of *total_bytes* and *value_count*? If that's true, it >>>>> seems that *total_bytes* would be more valuable for estimation >>>>> purposes since you have an explicit upper bound on size. >>>>> >>>>> -Dan >>>>> >>>>> On Tue, Sep 29, 2026 at 8:03 AM Eduard Tudenhöfner < >>>>> [email protected]> wrote: >>>>> >>>>>> Hey everyone, >>>>>> >>>>>> We had a few discussions around the *avg_value_size_in_bytes* field >>>>>> on the Iceberg slack and how it makes e.g. aggregations more difficult >>>>>> than >>>>>> necessary. We concluded that it's probably best to track the *total* >>>>>> instead of the *avg.* >>>>>> That being said, the field is being renamed to *total_bytes* in >>>>>> https://github.com/apache/iceberg/pull/18308. >>>>>> >>>>>> Please speak up if you have any concerns about this change. >>>>>> >>>>>> Thanks, >>>>>> Eduard >>>>>> >>>>>
