+1 for using total rather than average. On Tue, Sep 29, 2026 at 8:30 AM Steven Wu <[email protected]> wrote:
> Dan, > > Here is the slack thread: > https://apache-iceberg.slack.com/archives/C0BDHBAGARG/p1790113027624689 > > Yes, your assumption is correct. > > Thanks, > Steven > > > > > > On Tue, Sep 29, 2026 at 8:28 AM Daniel Weeks <[email protected]> wrote: > >> Eduard, >> >> Could you link to the slack discussion (I wasn't able to find it)? >> >> Is it safe to assume that we can determine the average via the >> combination of *total_bytes* and *value_count*? If that's true, it seems >> that *total_bytes* would be more valuable for estimation purposes since >> you have an explicit upper bound on size. >> >> -Dan >> >> On Tue, Sep 29, 2026 at 8:03 AM Eduard Tudenhöfner < >> [email protected]> wrote: >> >>> Hey everyone, >>> >>> We had a few discussions around the *avg_value_size_in_bytes* field on >>> the Iceberg slack and how it makes e.g. aggregations more difficult than >>> necessary. We concluded that it's probably best to track the *total* >>> instead of the *avg.* >>> That being said, the field is being renamed to *total_bytes* in >>> https://github.com/apache/iceberg/pull/18308. >>> >>> Please speak up if you have any concerns about this change. >>> >>> Thanks, >>> Eduard >>> >>
