claudevdm opened a new pull request, #40255:
URL: https://github.com/apache/beam/pull/40255

   * AddFiles: dry_run reports what the schema pre-pass would do without 
committing or registering
   
   SchemaEvolutionConfig.setDryRun(true) (provider key dry_run) turns AddFiles 
into a report-only transform: the read side runs as usual (footers, distinct 
schemas), then DryRunReport emits one Row per distinct schema plus one summary 
row on a new output, dry_run_report.
   
   Why: before enabling evolution on a large import, a user wants to know what 
the options would do to the table and which files would be refused, without 
touching anything. Running the real transform under FAIL_PIPELINE answers only 
"would it fail", and only for the first failure.
   
   The verdicts come from CommitSchemaUnion.plan, the one step that decides 
everything a commit does before writing anything. A Plan says which distinct 
file schemas are merged (SchemaToMerge), which are refused and why 
(IncompatibleSchema), and the schema the table ends with. It is an 
EvolutionPlan against an existing table (base schema snapshot, name-mapping 
repair) or a CreationPlan when the table is missing (partition spec and sort 
order resolved against the union, or the problem that blocks creation). 
commitOnce dispatches to evolve or create, which fail or warn under the 
handling mode and then write; the dry run turns the same plan into rows, so a 
check added to the plan reaches both and the report cannot drift from the 
commit. A schema that is fine against the table but conflicts with another 
schema of the input is therefore reported with the blame a real run assigns.
   
   Real-run changes that come with planning first: partition or sort fields 
that do not fit the union fail with a message naming them, before any catalog 
write and under either handling (the per-file fallback creation throws the same 
error, so it could never be routed); problems are reported before the 
transaction is opened; every transaction is checked against the one base 
snapshot; planning retries like committing; a window without schemas plans 
nothing. Settings (config, the handling resolved for the mode, 
NewTableSettings) replaces the three loose arguments both DoFns carried.
   
   Report rows (REPORT_SCHEMA), told apart by row_type:
     row_type         schema | create | unreadable | unchecked | summary
     schema_key       short murmur3 key of the schema JSON on schema rows
     schema           the canonical schema JSON; on the create row, the
                      union the table would be created with
     num_files        files the row covers
     changes          ARRAY<STRING>: the SchemaDelta descriptions on schema
                      rows; "create <optional|required> <name> <type>" per
                      column on the create row (pins shown required); the
                      totals line and table-level changes on the summary
     allowed          whether a real run would accept it
     reason           why not, else ""; the consequence on the summary
     would_create_table  false whenever a real run would not create the
                      table, including when it would fail first
   
   Provider: the dry_run_report output only exists when dry_run is set, so 
existing YAML pipelines that enumerate outputs are unaffected.
   
   * improve retry
   
   * comments
   
   * add untracked file
   
   * move tests
   
   * track file
   
   * comments
   
   **Please** add a meaningful description for your change here
   
   ------------------------
   
   Thank you for your contribution! Follow this checklist to help us 
incorporate your contribution quickly and easily:
   
    - [ ] Mention the appropriate issue in your description (for example: 
`addresses #123`), if applicable. This will automatically add a link to the 
pull request in the issue. If you would like the issue to automatically close 
on merging the pull request, comment `fixes #<ISSUE NUMBER>` instead.
    - [ ] Update `CHANGES.md` with noteworthy changes.
    - [ ] If this contribution is large, please file an Apache [Individual 
Contributor License Agreement](https://www.apache.org/licenses/icla.pdf).
   
   See the [Contributor Guide](https://beam.apache.org/contribute) for more 
tips on [how to make review process 
smoother](https://github.com/apache/beam/blob/master/CONTRIBUTING.md#make-the-reviewers-job-easier).
   
   To check the build health, please visit 
[https://github.com/apache/beam/blob/master/.test-infra/BUILD_STATUS.md](https://github.com/apache/beam/blob/master/.test-infra/BUILD_STATUS.md)
   
   GitHub Actions Tests Status (on master branch)
   
------------------------------------------------------------------------------------------------
   [![Build python source distribution and 
wheels](https://github.com/apache/beam/actions/workflows/build_wheels.yml/badge.svg?event=schedule&&?branch=master)](https://github.com/apache/beam/actions?query=workflow%3A%22Build+python+source+distribution+and+wheels%22+branch%3Amaster+event%3Aschedule)
   [![Python 
tests](https://github.com/apache/beam/actions/workflows/python_tests.yml/badge.svg?event=schedule&&?branch=master)](https://github.com/apache/beam/actions?query=workflow%3A%22Python+Tests%22+branch%3Amaster+event%3Aschedule)
   [![Java 
tests](https://github.com/apache/beam/actions/workflows/java_tests.yml/badge.svg?event=schedule&&?branch=master)](https://github.com/apache/beam/actions?query=workflow%3A%22Java+Tests%22+branch%3Amaster+event%3Aschedule)
   [![Go 
tests](https://github.com/apache/beam/actions/workflows/go_tests.yml/badge.svg?event=schedule&&?branch=master)](https://github.com/apache/beam/actions?query=workflow%3A%22Go+tests%22+branch%3Amaster+event%3Aschedule)
   
   See [CI.md](https://github.com/apache/beam/blob/master/CI.md) for more 
information about GitHub Actions CI or the [workflows 
README](https://github.com/apache/beam/blob/master/.github/workflows/README.md) 
to see a list of phrases to trigger workflows.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to