hutiefang76 opened a new pull request, #12498:
URL: https://github.com/apache/seatunnel/pull/12498

   ### Purpose of this pull request
   
   Closes #12497.
   
   With DuckDB JDBC 1.3.1, a 100-row `executeBatch` into DuckLake produced 100 
Parquet files in a local reproduction. The JDBC sink's `batch_size` did not 
prevent this. This PR adds an opt-in `ducklake_bulk_write` path that stages 
each batch in a connection-local DuckDB temporary table, then performs one 
`INSERT ... SELECT` into the attached DuckLake table. The ordinary DuckDB/JDBC 
path remains the default.
   
   ### Does this PR introduce any user-facing change?
   
   Yes. For an existing DuckLake table, users can set `generate_sink_sql = 
true`, `database = "lake"`, `table = "main.events"`, and `ducklake_bulk_write = 
true`. An attached lake can be initialized on each JDBC connection through 
DuckDB's `session_init_sql_file` URL parameter; the PostgreSQL database, 
metadata schema, and data path are configured separately in that SQL file. The 
sink's dry-run connection check validates the attached target and input columns 
without DDL/DML.
   
   The mode is append-only and requires `schema_save_mode = IGNORE`, 
`data_save_mode = APPEND_DATA`, `auto_commit = true`, `max_retries = 0`, and a 
positive `batch_size`. It rejects custom write SQL, primary keys/upserts, COPY, 
and XA. A task replay after an uncertain commit can still duplicate rows; this 
change does not claim exactly-once delivery. English and Chinese JDBC/DuckDB 
docs describe the setup and limits. The DuckDB docs also remove an XA example 
that referred to a class absent from the pinned 1.3.1 driver.
   
   ### How was this patch tested?
   
   - JDK 17: `mvn -o -pl seatunnel-connectors-v2/connector-jdbc -am -DskipTests 
-q verify` passed; module Spotless applied.
   - Targeted JDK 17 tests: `DuckLakeBulkWriteTest` (3), `JdbcSinkFactoryTest` 
(18), and existing `DuckDBSourceAndSinkTest` (1) all passed. The opt-in 
DuckLake test uses the matching DuckLake and SQLite scanner extension files, 
compares 100 files from ordinary JDBC batching with one additional file from 
the SeaTunnel bulk sink, runs the dry-run check, and reads back 100 distinct 
rows after reconnect. The other tests cover configuration rejection, an 
unsuccessful transfer followed by job-style replay, and 
decimal/timestamptz/null values through staging.
   - Manual deployment smoke: DuckDB JDBC 1.3.1 + matching extensions attached 
DuckLake to an isolated PostgreSQL `lake_metadata` database with a non-public 
`dl_meta` metadata schema and a MinIO S3 data path. A 100-row staged insert 
containing nullable text, decimal, and timestamptz values yielded one Parquet 
object and all 100 rows were readable after reconnect. The temporary containers 
were removed afterward.
   
   ### Check list
   
   * [x] No new Jar binary package or connector module is added.
   * [x] English and Chinese connector documentation is updated.
   * [x] The option defaults to false, so existing DuckDB writes are unchanged.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to