> I support "option 0," which I'll place at the top of the list as: > 'Changes in Parquet files are versioned in the current way.'
I think that the "current way" is that we don't make breaking changes. I voted against changing `path_in_schema` to optional because I think it breaks our guarantees. I think it's important for us to be able to agree on how to make breaking changes. On Sun, Sep 13, 2026 at 11:50 PM Will Edwards via dev < [email protected]> wrote: > Hi Ryan, thx for clarifying :thumb-up: > > So, I support "option 0," which I'll place at the top of the list as: > 'Changes in Parquet files are versioned in the current way.' > > Using path_in_schema as the example big braking change I could imagine > getting traction: > > When we make path-in-schema optional we might bump the version number of > the parquet-format package as a hint for implementers to read the release > notes, but we won't put a version gate on the files themselves. We expect > it to be opted-in only by those most affected by footer bloat and only on > those files most affected by that bloat. With the passage of time more > users might be reaching for it and the reader base will have matured and we > might contemplate some writers making it the default or a threshold-driven > default. But that'll take time. > > As you mentioned, Ryan, readers may not correctly follow Thrift versioning > and could be buggy or implementers might not read release notes. This is > true but I reason that those upgrading parquet-format without reading > release notes would also be those who don't know to gate on a new version > field, etc., in the other options on the list. Also, I know of one > mainstream widely-used Parquet reader that doesn't check the trailing file > magic for PAR1 either. So, it is what it is and the same problems will be > faced by any version-gate-in-file approach. > > When this was voted on previously I had thought the versioning discussion > would focus on making it easier to reason about what readers and writers > support in feature matrices and those kinds of decisions that users opting > into features want to know about, perhaps by bumping the version of > parquet-format. I wasn't thinking it'd be a version gate on the produced > Parquet files at rest themselves :D > > Best, > Will > > On Thu, 10 Sept 2026 at 21:25, Ryan Blue <[email protected]> wrote: > > > I'm also replying to Andrew here, but I think it helps to keep the > replies > > separate and focused on one topic. > > > > > From what I can tell, the current state of parquet is implicitly > Option 2 > > > (as readers can and do try to read any file and skip or error when > > > encountering unsupported features), and don't check version numbers. > > > > > [With option 2] evolving the format would have > > > additional requirements (but the same requirements already implicitly > > > exist) > > > > I don't agree that option 2 is the current state. I think the current > state > > is that we don't make these breaking changes, and that's the problem I > want > > to fix. > > > > We've released new encodings and compression that guarantee failure > because > > they use a new enum symbol, but will only break readers when those > columns > > are read. However, we have structured all of the other changes to be > > forward-compatible. For instance, when we found the sort bug for string > > columns, we introduced a second set of fields for lower and upper bounds. > > We also use structs to mimic enums when we want them to be forward > > compatible (like logical type annotations). > > > > Without a way to make and coordinate breaking changes, I think we must > > adopt a guarantee like the one in Option 2. If we do that, I think we've > > made it even harder to evolve the format and I don't see why we would > > decide to make a breaking change to fix something like path_in_schema. > > > > There's also another way to look at this: if we are confident that > > path_in_schema will break all older readers, why not use that > compatibility > > break to get other cleanup features in? Doing that is one of the > advantages > > of bundling. > > > > On Thu, Sep 10, 2026 at 2:12 AM Andrew Lamb <[email protected]> > > wrote: > > > > > Here is my attempt to summarize the tradeoffs (Ryan's explanation > during > > > the call was very helpful for me). > > > > > > The core tradeoff is in requirements for future changes to the Parquet > > spec > > > itself vs how many files particular readers can read. > > > > > > If the spec requires that readers fail fast on unknown versions > ("Option > > > 1"), > > > * Pro: New changes to the spec don't have to consider existing readers, > > and > > > are thus in theory are easier/faster to make (e.g. Ryan's example of > > > relocatable page headers) > > > * Con: requires readers to fail on all files with newer versions, even > > > those files that the reader could have read correctly > > > > > > If the spec allows readers to attempt to read unknown versions ("Option > > 2") > > > * Pro: Readers will be able to read more files, though it will be > harder > > to > > > reason up front if a reader can read any file (may have to test it) > > > * Con: There is a requirement on any future changes to the spec to > ensure > > > old readers don't interpret new features ("additional guarantee that > > > reading future formats will either fail or produce correct results"). > > This > > > requirement is hard to define precisely given the wide and unknown > > variety > > > of readers > > > > > > From what I can tell, the current state of parquet is implicitly > Option 2 > > > (as readers can and do try to read any file and skip or error when > > > encountering unsupported features), and don't check version numbers. > > > > > > My personal opinion is that allowing readers to read unknown versions > > > (option 2 / option 3) is the most practical: Changing the current > > implicit > > > behavior would be quite confusing, and making the spec harder to change > > for > > > wider read interoperability is the right tradeoff in my mind. > > > > > > Andrew > > > > > > > > > p.s. > > > > > > > I think path_in_schema is a good example of why option 2 will > > inevitably > > > produce > > > correctness bugs > > > > > > It seems to me that making path_in_schema optional will simply cause > old > > > readers to fail if they need it, so is not a good example of why > option 2 > > > would necessarily correctness issues compared to option 1. The example > > > about java hash sets could be avoided with adequate testing, for > example, > > > and I don't see how gating that code behind reading a new version > number > > > would make that bug any more/less likely. > > > > > > > My second argument for why we should not attempt to read all future > > > versions > > > of Parquet is that it ends up limiting how we can evolve the format. > > > > > > This makes a lot of sense to me -- evolving the format would have > > > additional requirements (but the same requirements already implicitly > > > exist); Adding relocatable pages, for example could be achieved by > > adding a > > > new DataPageHeaderV3 rather than modifying the existing structure. That > > is > > > more complicated to be sure, but not impossible. > > > > > > > > > > > > > > > On Thu, Sep 10, 2026 at 4:39 AM Antoine Pitrou <[email protected]> > > wrote: > > > > > > > Le 10/09/2026 à 00:38, Ryan Blue a écrit : > > > > > > > > > > There are 3 main options: > > > > > 1. A reader should fail because it does not support the version > > > > > 2. A reader should attempt to read the file > > > > > 3. This choice is left up to implementations > > > > > > > > > > I'll cover each option in more detail below, but first I want to > > > clarify > > > > > that we are not talking about "preview" features like encodings or > > > > > forward-compatible changes like new logical types. Preview features > > > will > > > > > break readers that do not support them and only affect specific > > columns > > > > > using the feature. For preview features, the expectation is that > > > readers > > > > > will attempt to read the file and will fail if they need to > project a > > > > > column that cannot be read. > > > > > > > > "Preview features" is a very weird terminology. It sounds like > > > > "unfinished" or "experimental". > > > > > > > > > The choice of how to handle an unsupported format version primarily > > > > affects > > > > > changes that add, remove, or modify the semantics of metadata > fields. > > > For > > > > > example: > > > > > - Changing `path_in_schema` from required to optional > > > > > > > > This depends whether the Thrift parser checks that required fields > are > > > > actually present in the serialized payload? Do we know what their > > > > current behavior is? > > > > > > > > i.e., does a Thrift parser generated with a required `path_in_schema` > > > > specification accept a serialized payload without that field? > > > > > > > > It also depends what the Parquet reader actually *does* with the > > > > `path_in_schema`? AFAICT, the Parquet C++ reader isn't doing anything > > > > specific with it. > > > > > > > > > Option 3 would mean that readers may choose to attempt to read, but > > do > > > > not > > > > > have additional guarantees. > > > > > > > > Option 3 can also mean "the reader is exposing an option to let the > > user > > > > choose the behavior (reject up front or attempt to read anyway)". > > > > > > > > (this would be more costly to implement, so I'm not sure any > > > > implementation would actually do that; but it's at least conceptually > > > > possible) > > > > > > > > Regards > > > > > > > > Antoine. > > > > > > > > > > > > > > > > > >
