zhangshenghang opened a new issue, #12524:
URL: https://github.com/apache/seatunnel/issues/12524

   ### Description
   
   For OceanBase in Oracle compatible mode (`jdbc:oceanbase:` + 
`compatible_mode = oracle`) the JDBC source uses the plain Oracle dialect. The 
OceanBase JDBC driver by default does not stream result sets server-side: 
without `useServerPrepStmts=true` the driver loads the complete result set into 
the client JVM, regardless of the configured fetch size. Reading a large table 
this way leads to heavy memory usage / OOM, and the sampling-based sharding in 
the chunk splitter has the same buffering problem because it reads a large 
sampled result set at once.
   
   Expected:
   1. Source connections for OceanBase Oracle mode should force 
`useServerPrepStmts=true` so rows are streamed by fetch size (sink connections 
keep user settings).
   2. Sampling-based sharding should be disabled for this dialect (fall back to 
bounded chunk-boundary queries), since the sampled result set may be buffered 
in memory by the driver.
   
   ### Impact
   
   JDBC source jobs reading large OceanBase (Oracle mode) tables risk OOM even 
with a small fetch size configured.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to