Russell's analysis makes sense. +1 for file-level flag. On Tue, Oct 6, 2026 at 6:08 PM Szehon Ho <[email protected]> wrote:
> Thanks Anoop for the clear writeup for the problem, makes sense to me for > per-file tightness for simplicity. > > The above list of use cases is an interesting food for thought, from > the analysis it sounds like most will repair all bounds and not need per > column after all, hence agree. > > Thanks > Szehon > > On Tue, Oct 6, 2026 at 2:58 PM Hongyue Zhang <[email protected]> > wrote: > >> +1 on the file-wide flag. >> >> From what I can tell, whether a given lower/upper bound is tight comes >> down to two parts. >> >> The first is file-level and changes over time. A DV added later deletes >> some rows, and the accuracy of the original bounds are no longer >> guaranteed. The new file wide flag can help record the change. >> >> The second is column-level and static. Variable-length columns like >> String and Binary can have truncated bounds, so we cannot make them >> reliably tight, and we don't have a use case that needs their tightness >> today. That's a property of the column type in the schema, so it doesn't >> need additional metadata for recording. >> >> So moving the tightness boolean from per-column to per-entry >> (tracked_file in v4) simplifies the bookkeeping for writers without giving >> up anything we care about. >> >> One note on upgrading the existing tables to v4. Since we want that to >> stay a lightweight metadata operation with no manifest rewrite, tightness >> is simply undefined for existing entries. Today the merge-on-read table >> does not have tight bounds. Once position deletes and DVs are rewritten to >> be colocated with the data file in a single tracked-file entry, tightness >> becomes a well-defined property of that entry, and a writer at that point >> can set the flag accordingly. >> >> Thanks, >> Hongyue >> >> On Tue, Oct 6, 2026 at 12:23 AM Eduard Tudenhöfner < >> [email protected]> wrote: >> >>> The file-wide flag seems to make sense to me, so +1 for that >>> >>> On Mon, Oct 5, 2026 at 10:22 PM Daniel Weeks <[email protected]> wrote: >>> >>>> +1 to Russell's analysis here. >>>> >>>> While we can construct hypothetical use cases where there may be >>>> benefit, it seems like those use cases are unlikely and/or narrow. >>>> >>>> I'd also suggest going with the file-wide flag. >>>> >>>> -Dan >>>> >>>> On Mon, Oct 5, 2026 at 12:07 PM Russell Spitzer < >>>> [email protected]> wrote: >>>> >>>>> I've been thinking through the patterns where per-column tightness >>>>> would matter. >>>>> >>>>> 1. >>>>> >>>>> *Selectively marking variable-length columns tight*. I don't have >>>>> a use case for this yet. A file-level flag can still mean every >>>>> non-truncated column is tight. >>>>> 2. >>>>> >>>>> *Selectively re-tightening a few columns, for example after a >>>>> column update*. Dan's point from the discussion applies here: a >>>>> workload whose DVs invalidate stats will also invalidate new column >>>>> stats. >>>>> A table left in a state where only some stats are tight due to an >>>>> update >>>>> seems unlikely. >>>>> 3. >>>>> >>>>> *A DV writer that can write tight stats*. This is the pattern that >>>>> matters to me. The writer repairs bounds while it deletes, and it can >>>>> repair every column it tracks. We still need an explicit flag, because >>>>> otherwise readers treat any attached DV as wide and ignore the repaired >>>>> bounds. One file-level flag covers this. >>>>> 4. >>>>> >>>>> *A command that restores tight stats when the DV writer cannot.* >>>>> This is worth supporting when deletes are infrequent: write the DVs, >>>>> analyze once, and use the tight bounds until the next DV. If we >>>>> frequently >>>>> write DV's then we are in the same bad situation as #2. Either way the >>>>> restore is all-or-nothing. I thought for a bit that there could be a >>>>> command that only restores stats for select columns, this could make >>>>> sense >>>>> with column update files but also feels like unecessary complexity. >>>>> >>>>> Patterns 3 and 4 are the realistic ones, and both only need the >>>>> file-wide flag. So I'm kind of leaning in that direction now. >>>>> >>>>> On Mon, Oct 5, 2026 at 12:32 PM Anoop Johnson <[email protected]> >>>>> wrote: >>>>> >>>>>> Hi, everyone - >>>>>> >>>>>> In v4, we introduced the notion of *tightness* in stats. Tightness >>>>>> indicates whether the upper/lower bounds truly exist in the live rows of >>>>>> a >>>>>> file. If the stats are tight, then engines can correctly answer many >>>>>> queries purely by consulting the metadata. >>>>>> >>>>>> However, when a DV is attached to a file, the stats are assumed to be >>>>>> non-tight, so these metadata-only query optimizations won't work >>>>>> anymore. >>>>>> >>>>>> So we want a way to represent stats tightness even in the presence of >>>>>> a DV. I have a short writeup >>>>>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0> >>>>>> of the problem and a few possible solutions. We discussed this today >>>>>> at the v4 adaptive metadata tree sync, and I wanted to continue the >>>>>> discussion. >>>>>> >>>>>> Please take a look at the doc >>>>>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0#heading=h.9cvri5ibsn4r> >>>>>> and let us know what you think by leaving a comment in the doc or in this >>>>>> email thread. >>>>>> >>>>>> Best, >>>>>> Anoop >>>>>> >>>>>
