Thanks everyone! I think we are converging on a file-level tightness flag. About the upgrade scenario Hongyue brought up - yes, files living inside existing v3 manifests won't have the tightness flag. In the absence of tightness, we assume that the stats are wide. New files created after the v4 upgrade should have it - either a `true` if there are no DVs, which means stats are tight, or a `false` if there is an attached DV, which means stats are wide.
On Thu, Oct 8, 2026 at 10:45 AM Steven Wu <[email protected]> wrote: > Russell's analysis makes sense. +1 for file-level flag. > > On Tue, Oct 6, 2026 at 6:08 PM Szehon Ho <[email protected]> wrote: > >> Thanks Anoop for the clear writeup for the problem, makes sense to me for >> per-file tightness for simplicity. >> >> The above list of use cases is an interesting food for thought, from >> the analysis it sounds like most will repair all bounds and not need per >> column after all, hence agree. >> >> Thanks >> Szehon >> >> On Tue, Oct 6, 2026 at 2:58 PM Hongyue Zhang <[email protected]> >> wrote: >> >>> +1 on the file-wide flag. >>> >>> From what I can tell, whether a given lower/upper bound is tight comes >>> down to two parts. >>> >>> The first is file-level and changes over time. A DV added later deletes >>> some rows, and the accuracy of the original bounds are no longer >>> guaranteed. The new file wide flag can help record the change. >>> >>> The second is column-level and static. Variable-length columns like >>> String and Binary can have truncated bounds, so we cannot make them >>> reliably tight, and we don't have a use case that needs their tightness >>> today. That's a property of the column type in the schema, so it doesn't >>> need additional metadata for recording. >>> >>> So moving the tightness boolean from per-column to per-entry >>> (tracked_file in v4) simplifies the bookkeeping for writers without giving >>> up anything we care about. >>> >>> One note on upgrading the existing tables to v4. Since we want that to >>> stay a lightweight metadata operation with no manifest rewrite, tightness >>> is simply undefined for existing entries. Today the merge-on-read table >>> does not have tight bounds. Once position deletes and DVs are rewritten to >>> be colocated with the data file in a single tracked-file entry, tightness >>> becomes a well-defined property of that entry, and a writer at that point >>> can set the flag accordingly. >>> >>> Thanks, >>> Hongyue >>> >>> On Tue, Oct 6, 2026 at 12:23 AM Eduard Tudenhöfner < >>> [email protected]> wrote: >>> >>>> The file-wide flag seems to make sense to me, so +1 for that >>>> >>>> On Mon, Oct 5, 2026 at 10:22 PM Daniel Weeks <[email protected]> wrote: >>>> >>>>> +1 to Russell's analysis here. >>>>> >>>>> While we can construct hypothetical use cases where there may be >>>>> benefit, it seems like those use cases are unlikely and/or narrow. >>>>> >>>>> I'd also suggest going with the file-wide flag. >>>>> >>>>> -Dan >>>>> >>>>> On Mon, Oct 5, 2026 at 12:07 PM Russell Spitzer < >>>>> [email protected]> wrote: >>>>> >>>>>> I've been thinking through the patterns where per-column tightness >>>>>> would matter. >>>>>> >>>>>> 1. >>>>>> >>>>>> *Selectively marking variable-length columns tight*. I don't have >>>>>> a use case for this yet. A file-level flag can still mean every >>>>>> non-truncated column is tight. >>>>>> 2. >>>>>> >>>>>> *Selectively re-tightening a few columns, for example after a >>>>>> column update*. Dan's point from the discussion applies here: a >>>>>> workload whose DVs invalidate stats will also invalidate new column >>>>>> stats. >>>>>> A table left in a state where only some stats are tight due to an >>>>>> update >>>>>> seems unlikely. >>>>>> 3. >>>>>> >>>>>> *A DV writer that can write tight stats*. This is the pattern >>>>>> that matters to me. The writer repairs bounds while it deletes, and >>>>>> it can >>>>>> repair every column it tracks. We still need an explicit flag, because >>>>>> otherwise readers treat any attached DV as wide and ignore the >>>>>> repaired >>>>>> bounds. One file-level flag covers this. >>>>>> 4. >>>>>> >>>>>> *A command that restores tight stats when the DV writer cannot.* >>>>>> This is worth supporting when deletes are infrequent: write the DVs, >>>>>> analyze once, and use the tight bounds until the next DV. If we >>>>>> frequently >>>>>> write DV's then we are in the same bad situation as #2. Either way the >>>>>> restore is all-or-nothing. I thought for a bit that there could be a >>>>>> command that only restores stats for select columns, this could make >>>>>> sense >>>>>> with column update files but also feels like unecessary complexity. >>>>>> >>>>>> Patterns 3 and 4 are the realistic ones, and both only need the >>>>>> file-wide flag. So I'm kind of leaning in that direction now. >>>>>> >>>>>> On Mon, Oct 5, 2026 at 12:32 PM Anoop Johnson <[email protected]> >>>>>> wrote: >>>>>> >>>>>>> Hi, everyone - >>>>>>> >>>>>>> In v4, we introduced the notion of *tightness* in stats. Tightness >>>>>>> indicates whether the upper/lower bounds truly exist in the live rows >>>>>>> of a >>>>>>> file. If the stats are tight, then engines can correctly answer many >>>>>>> queries purely by consulting the metadata. >>>>>>> >>>>>>> However, when a DV is attached to a file, the stats are assumed to >>>>>>> be non-tight, so these metadata-only query optimizations won't work >>>>>>> anymore. >>>>>>> >>>>>>> So we want a way to represent stats tightness even in the presence >>>>>>> of a DV. I have a short writeup >>>>>>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0> >>>>>>> of the problem and a few possible solutions. We discussed this >>>>>>> today at the v4 adaptive metadata tree sync, and I wanted to continue >>>>>>> the >>>>>>> discussion. >>>>>>> >>>>>>> Please take a look at the doc >>>>>>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0#heading=h.9cvri5ibsn4r> >>>>>>> and let us know what you think by leaving a comment in the doc or in >>>>>>> this >>>>>>> email thread. >>>>>>> >>>>>>> Best, >>>>>>> Anoop >>>>>>> >>>>>>
