Coincidentally I ran into this as well while working on the evaluator for
content stats (see
https://github.com/apache/iceberg/pull/17413/changes/53456855fbf126502a7e6a15ce39e1833f40d43f
from
https://github.com/apache/iceberg/pull/17413) but didn't have time yet to
discuss this issue with the community. I think that including the
*null_value_count* for required fields inside optional parents makes sense.

On Mon, Aug 10, 2026 at 11:41 PM Xiening Dai <[email protected]> wrote:

> Hi Anoop,
>
> I like your "effective nullability" definition. The current spec wording
> definitely needs update:
>
> ```
> A struct is a tuple of typed values. Each field in the tuple is named and
> has an integer id that is unique in the table schema. Each field can be
> either optional or required, meaning that values can (or cannot) be null.
> Fields may be any type.
> ```
>
> And I agree we should remove the "optional null_value_count" comment, and
> writer should record the null count whenever possible for nested fields
> (regardless required or optional). Checking the null count against 0 is
> much easier than going through and checking the ancestors nullability.
>
> On 2026/08/10 20:09:00 Anoop Johnson wrote:
> > Hi, Xiening -
> >
> > What you brought up is an important subtlety. I view it a bit differently
> > though:  what matters is *effective* nullability along the whole path
> from
> > the root, not the leaf field's own *required* flag. In your
> > specific example,  `s.a` is declared NOT NULL, but its parent `s` is
> > optional. So effectively `s.a` is nullable.   A field is only truly
> > non-nullable if it is required and every ancestor struct up to the root
> is
> > also required. A genuinely non-nullable field like this does not require
> > `null_value_count` (because it will always be 0).
> >
> > So `s.a` here is not a required field in the spec perspective, and
> > null_value_count *should* be recorded for it.  I believe the current
> > Iceberg stats collector and evaluator handles this correctly. e.g. the
> > InclusiveMetricsEvaluator relies
> > <
> https://github.com/apache/iceberg/blob/89e2f887491c1b5fa9f8b9de81b3aa8b31fa6974/api/src/main/java/org/apache/iceberg/expressions/InclusiveMetricsEvaluator.java#L100-L102
> >
> > on the explicit count and not the field-level required flag.
> >
> >   @Override
> >     protected boolean mayContainNull(int id) {
> >       return nullCounts == null || !nullCounts.containsKey(id) ||
> > nullCounts.get(id) != 0;
> >    }
> >
> > Overall, I agree with the spirit of your suggestion.  We could clarify
> the
> > spec wording so "optional fields" reads as *effectively required*, and
> make
> > explicit that a field is non-nullable only when it and all of its
> ancestors
> > are required.
> >
> > Best,
> > Anoop
> >
> > On Mon, Aug 10, 2026 at 10:45 AM Xiening Dai <[email protected]> wrote:
> >
> > > Hi all,
> > >
> > > I see this in V4 spec: "null_value_count is only used for optional
> fields"
> > > (
> > >
> https://github.com/apache/iceberg/blob/89e2f887491c1b5fa9f8b9de81b3aa8b31fa6974/format/spec.md?plain=1#L838
> > > )
> > >
> > > This only makes sense if required field never contains nulls. But that
> > > actually is not true. In table schema, we can have a struct field
> which is
> > > optional and has a child field that is required. In that case, when the
> > > parent struct is null, the required field values is treated as NULL
> too.
> > >
> > > For example, below works in Spark:
> > >
> > > CREATE TABLE t (i INT, s STRUCT<a: INT NOT NULL, b: STRING>) USING
> iceberg;
> > > INSERT INTO t VALUES (1, NULL);
> > > SELECT COUNT (*) WHERE s.a IS NULL; -- returns 1
> > >
> > > Because of this, I don't think the query engine can optimize the
> execution
> > > plan and assume a field doesn't contain null when it's marked as
> required.
> > > It would need to rely on the null_value_count explicitly being 0. If
> you
> > > agree on this, then we should always record the null_value_count no
> matter
> > > if it's required or optionals field.
> > >
> > > Thoughts?
> > >
> >
>

Reply via email to