zhangshenghang opened a new pull request, #12530:
URL: https://github.com/apache/seatunnel/pull/12530

   ### Purpose of this pull request
   
   Close #12524.
   
   For OceanBase in Oracle compatible mode (`jdbc:oceanbase:` + 
`compatible_mode = oracle`), the JDBC source used the plain `OracleDialect`. 
The OceanBase JDBC driver by default does not stream result sets server-side: 
without `useServerPrepStmts=true` the driver loads the complete result set into 
the client JVM regardless of the configured fetch size, so reading a large 
table can lead to OOM. Sampling-based sharding has the same problem because it 
reads a large sampled result set at once.
   
   This PR:
   
   1. Adds `OceanBaseOracleDialect` (returned by `OceanBaseDialectFactory` for 
Oracle mode). It extends `OracleDialect` and keeps the `Oracle` dialect name, 
because Oracle-specific execution paths are keyed on the dialect name and 
OceanBase Oracle mode is Oracle compatible - so no existing behavior changes.
   2. Adds a `configureSourceConnection` hook on `JdbcDialect` (used by 
`JdbcSourceFactory`), which the new dialect overrides to force 
`useServerPrepStmts=true` on source connections so rows are streamed by fetch 
size. Sink connections keep the user's settings.
   3. Adds a `supportsSamplingSharding` hook on `JdbcDialect` and disables 
sampling-based sharding for OceanBase Oracle mode, falling back to bounded 
chunk-boundary queries. The existing `split.sample-sharding.*` user options 
still apply on top of this for all other dialects.
   
   ### Does this PR introduce _any_ user-facing change?
   
   Yes, for OceanBase Oracle mode sources only: `useServerPrepStmts=true` is 
now forced on source connections (a warning is logged when the user had 
configured it as `false`), and sampling-based sharding is skipped in favor of 
bounded chunk splitting. Both changes prevent the driver from buffering the 
complete result set in memory. All other dialects and the sink are unchanged.
   
   ### How was this patch tested?
   
   - Added `OceanBaseOracleDialectTest`: only OceanBase Oracle disables 
sampling sharding; source connection forces the server-side cursor (also 
verified through the OceanBase driver `UrlParser`); sink connection keeps user 
settings; factory wiring returns the new dialect with the `Oracle` dialect name.
   - Extended `DynamicChunkSplitterTest`: OceanBase Oracle skips sampling 
sharding; other dialects follow the user option.
   - `mvn -pl seatunnel-connectors-v2/connector-jdbc test`: all tests passing.
   - No E2E added: there is no OceanBase Oracle E2E environment in the 
repository; the connection-parameter behavior is covered by the driver-level 
`UrlParser` unit test.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to