Thanks for the summary Rok, I like the direction C1 is going, but I'm a little concerned about the following statement:
Regarding NaN/Inf limitations, the consensus was that these should be > handled by statistics. This leads to a need for a new inf_count statustics. I feel like the intent behind having a logical type for vector is to differentiate its semantics from a simple fixed size list. Vectors shouldn't contain null, NaN, inf as elements (in practice they may present, but those are typically due to erroneous handling). I would expect vectors to prohibit special/missing values that render them invalid. Are the statistics intended to detect these erroneous scenarios or is the intent to allow these values? -Dan On Tue, Sep 8, 2026 at 1:41 PM Will Edwards via dev <[email protected]> wrote: > Hi all, > > Andrew's and Gunnar's results show that existing LIST storage can be read > efficiently. This matches my own experience too. I think efficiency is a > software problem not a Parquet format problem. E.g. the reader can expose > flat typed memory; nothing in the format requires allocating a List<Float> > per row. A nullable schema doesn't change that for pages without nulls, > which the reader can establish at page scope. > > So is the immediate question what SHOW CREATE TABLE should say when the > only source of schema is a Parquet file? ARRAY<FLOAT>, a fixed-length > array, or VECTOR<FLOAT, 768>? That's a useful discussion about the contract > we're expressing. > > SELECT COS_SIM(a, b) etc can already work on numeric lists, with lengths > checked efficiently through the encoded levels. Whether a and b belong to > the same embedding space remains the user's responsibility; declaring both > to have 768 dimensions doesn't establish that. > > There are other semantics I'd be interested in discussing: > > - Vector indexes, along the lines of Lance: how does a reader discover > that a column or auxiliary structure indexes another column, or is derived > from it? What lets an engine recognise when it can use that structure? > > - Quantisation, normalisation guarantees, and additional numeric formats > such as BF16/FP8. We already have FLOAT16 and small integer annotations, > but a coordinate type alone doesn't describe all of these. Some quantised > representations are naturally composite values: packed coordinates plus > scales, offsets, or codebooks. Block-scaled FP4 is a concrete example. > These could be structs with separate children for codes and scales, and a > logical annotation describing how they collectively represent a vector. > Shared parameters might live elsewhere. Either way, the relationship needs > expressing. > > Should we pin down those requirements before choosing a physical > representation? Adding an element-type enum doesn't answer all of these, > and packing whole vectors into FLBA limits existing element encoding > choices. > > What concrete vector workload still needs a physical change after applying > the reader optimisations we've discussed? > > Will > > On Tue, 8 Sept 2026 at 21:22, Antoine Pitrou <[email protected]> wrote: > > > > > Le 08/09/2026 à 20:42, Rok Mihevc a écrit : > > > Hi all, > > > > > > As proposed during the general community call we held a focused call > > today > > > to discuss the proposal for adding a logical vector type to parquet. > This > > > was called because it seems: > > > > > > * Vectors (as in vector DB) storage needs are different to what we need > > for > > > fixed size list applications we've been disucssing so far [0]. > > > > This seems unsubstantiated. Can you please provide an explanation? > > > > Ideally we'd use the *same* underlying solution for both problems. > > > > > We have converged on two options that could work, but have not yet > > decided > > > which one to propose: > > > > > > - Option A - FLBA as physical storage (FLBA(num_elements * > > element_width)). > > > Vector logical type would provide an enum specifying the element type, > > and > > > the element would be non-nullable. The element type enum could be > > > independent of the parquet type system, making evolution easier. > Downside > > > is loss of element-level encodings available today, though they can be > > > added later. Upside is simplicity and ease of evolution. > > > > > > - Option C1 - LIST with numeric parquet type elements as physical > > storage. > > > New float types (e.g. BF16) can be added in parquet proper as needed. > > > Downside is we need to add new types to parquet, DL/RL overhead is > likely > > > not a concern. Upside is we have encoding support now. > > > > Option A is a short-term solution with severe limitations (lack of > > encoding choices) that make me entirely negative about the idea. > > > > I'm lukewarm about Option C (especially the optimistic statement that > > "DL/RL overhead is likely not a concern"), but at least it doesn't paint > > us in a corner. > > > > Regards > > > > Antoine. > > > > > > >
