Russell's analysis makes sense. +1 for file-level flag.

On Tue, Oct 6, 2026 at 6:08 PM Szehon Ho <[email protected]> wrote:

> Thanks Anoop for the clear writeup for the problem, makes sense to me for
> per-file tightness for simplicity.
>
> The above list of use cases is an interesting food for thought, from
> the analysis it sounds like most will repair all bounds and not need per
> column after all, hence agree.
>
> Thanks
> Szehon
>
> On Tue, Oct 6, 2026 at 2:58 PM Hongyue Zhang <[email protected]>
> wrote:
>
>> +1 on the file-wide flag.
>>
>> From what I can tell, whether a given lower/upper bound is tight comes
>> down to two parts.
>>
>> The first is file-level and changes over time. A DV added later deletes
>> some rows, and the accuracy of the original bounds are no longer
>> guaranteed. The new file wide flag can help record the change.
>>
>> The second is column-level and static. Variable-length columns like
>> String and Binary can have truncated bounds, so we cannot make them
>> reliably tight, and we don't have a use case that needs their tightness
>> today. That's a property of the column type in the schema, so it doesn't
>> need additional metadata for recording.
>>
>> So moving the tightness boolean from per-column to per-entry
>> (tracked_file in v4) simplifies the bookkeeping for writers without giving
>> up anything we care about.
>>
>> One note on upgrading the existing tables to v4. Since we want that to
>> stay a lightweight metadata operation with no manifest rewrite, tightness
>> is simply undefined for existing entries. Today the merge-on-read table
>> does not have tight bounds. Once position deletes and DVs are rewritten to
>> be colocated with the data file in a single tracked-file entry, tightness
>> becomes a well-defined property of that entry, and a writer at that point
>> can set the flag accordingly.
>>
>> Thanks,
>> Hongyue
>>
>> On Tue, Oct 6, 2026 at 12:23 AM Eduard Tudenhöfner <
>> [email protected]> wrote:
>>
>>> The file-wide flag seems to make sense to me, so +1 for that
>>>
>>> On Mon, Oct 5, 2026 at 10:22 PM Daniel Weeks <[email protected]> wrote:
>>>
>>>> +1 to Russell's analysis here.
>>>>
>>>> While we can construct hypothetical use cases where there may be
>>>> benefit, it seems like those use cases are unlikely and/or narrow.
>>>>
>>>> I'd also suggest going with the file-wide flag.
>>>>
>>>> -Dan
>>>>
>>>> On Mon, Oct 5, 2026 at 12:07 PM Russell Spitzer <
>>>> [email protected]> wrote:
>>>>
>>>>> I've been thinking through the patterns where per-column tightness
>>>>> would matter.
>>>>>
>>>>>    1.
>>>>>
>>>>>    *Selectively marking variable-length columns tight*. I don't have
>>>>>    a use case for this yet. A file-level flag can still mean every
>>>>>    non-truncated column is tight.
>>>>>    2.
>>>>>
>>>>>    *Selectively re-tightening a few columns, for example after a
>>>>>    column update*. Dan's point from the discussion applies here: a
>>>>>    workload whose DVs invalidate stats will also invalidate new column 
>>>>> stats.
>>>>>    A table left in a state where only some stats are tight due to an 
>>>>> update
>>>>>    seems unlikely.
>>>>>    3.
>>>>>
>>>>>    *A DV writer that can write tight stats*. This is the pattern that
>>>>>    matters to me. The writer repairs bounds while it deletes, and it can
>>>>>    repair every column it tracks. We still need an explicit flag, because
>>>>>    otherwise readers treat any attached DV as wide and ignore the repaired
>>>>>    bounds. One file-level flag covers this.
>>>>>    4.
>>>>>
>>>>>    *A command that restores tight stats when the DV writer cannot.*
>>>>>    This is worth supporting when deletes are infrequent: write the DVs,
>>>>>    analyze once, and use the tight bounds until the next DV. If we 
>>>>> frequently
>>>>>    write DV's then we are in the same bad situation as #2. Either way the
>>>>>    restore is all-or-nothing. I thought for a bit that there could be a
>>>>>    command that only restores stats for select columns, this could make 
>>>>> sense
>>>>>    with column update files but also feels like unecessary complexity.
>>>>>
>>>>> Patterns 3 and 4 are the realistic ones, and both only need the
>>>>> file-wide flag. So I'm kind of leaning in that direction now.
>>>>>
>>>>> On Mon, Oct 5, 2026 at 12:32 PM Anoop Johnson <[email protected]>
>>>>> wrote:
>>>>>
>>>>>> Hi, everyone -
>>>>>>
>>>>>> In v4, we introduced the notion of *tightness* in stats. Tightness
>>>>>> indicates whether the upper/lower bounds truly exist in the live rows of 
>>>>>> a
>>>>>> file. If the stats are tight, then engines can correctly answer many
>>>>>> queries purely by consulting the metadata.
>>>>>>
>>>>>> However, when a DV is attached to a file, the stats are assumed to be
>>>>>> non-tight, so these  metadata-only query optimizations won't work 
>>>>>> anymore.
>>>>>>
>>>>>> So we want a way to represent stats tightness even in the presence of
>>>>>> a DV. I have a short writeup
>>>>>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0>
>>>>>> of the problem and a few possible solutions. We discussed this today
>>>>>> at the v4 adaptive metadata tree sync, and I wanted to continue the
>>>>>> discussion.
>>>>>>
>>>>>> Please take a look at the doc
>>>>>> <https://docs.google.com/document/d/13DbwnwKQXjvY3vRcFJo-qVAQWk-GB4zLzzYmFdPntZg/edit?tab=t.0#heading=h.9cvri5ibsn4r>
>>>>>> and let us know what you think by leaving a comment in the doc or in this
>>>>>> email thread.
>>>>>>
>>>>>> Best,
>>>>>> Anoop
>>>>>>
>>>>>

Reply via email to