Engines that disallow nulls, nans and infs: - postgress - oracle - mysql - sql server - also: milvus, qdrant, pinecone
Notable dbs that allow nulls, nans and infs: - duckdb - clickhouse These seem to have this behavior because they do not have a specific vector type so they use array(float) or fixed-size-array(float) and disable functionality when nulls or non-finite values are present. So VECTOR across most database engines disallow nulls, nans and infs. This means better interoperability between parquet and such dbs so I think we should move ahead with the proposal as is. On Wed, Sep 23, 2026 at 5:00 PM Antoine Pitrou <[email protected]> wrote: > > Le 23/09/2026 à 16:26, Rok Mihevc a écrit : > >> Well, Parquet is not a single-purpose format. > >> > >> Vector stores can choose whatever policy fits their usage, but we're > >> talking about a general-purpose file format that aims to be broadly > >> applicable. > > > > What is the gate for Parquet to define something as a logical type? > > There is no "gate". But we should act in the broader community's > interest. A limitation that is only motivated by the needs of a single > segment of the user base goes against that goal. > > >> Why not? Why shouldn't it be used for storing, for example, NumPy tensor > >> data (which can contain NaNs and infinites)? > > > > A multidimensional arrays logical type might come with other > > requirements (shapes, strides, dimension names), that vector > > applications might not have. > > Well, first, you can use NumPy just for 1D data, and I'm sure some > people do that. Second, strides wouldn't be saved in Parquet where they > are meaningless. Third, all of this can be conveyed as additional > key-value metadata. > > >> It does not make sense to have *both* a "specialist" Vector type and a > >> "general-purpose" FixedSizeList type. Parquet types are a > >> general-purpose vocabulary from which you can build up more specialized > >> applications. Specialist types can be left to domain-specific SW users > >> of Parquet. > >> > >> (for example, Arrow implementations serialize the Arrow schema in > >> Parquet metadata so that the Arrow schema can be rebuilt at read time, > >> without mandating that every Arrow datatype has a Parquet equivalent) > > > > That does seem to work nicely for Arrow [0]. Should we revisit the > Parquet > > extensions discussion [1][2]? > > I think we should, and then specialized communities such as ML can > define their own specialized extension types (with arbitrary semantics) > on top of regular Parquet logical/physical types. > > Regards > > Antoine. > > >
