zhangshenghang opened a new pull request, #12159:
URL: https://github.com/apache/seatunnel/pull/12159

   ## Purpose
   
   Some JDBC dialects return a zero or negative estimate from 
`queryApproximateRowCnt`:
   
   - AnalyticDB MySQL external tables, which do not have accurate table 
statistics.
   - MySQL tables whose `information_schema` statistics have not been updated.
   - PostgreSQL tables that have never been `ANALYZE`d.
   
   With `approximateRowCnt <= 0` the previous code path went through the 
uneven-distribution branch and `splitUnevenlySizedChunks`, which issues one 
extra boundary query per chunk. For large uneven tables this is slow and 
produces a noisy split log.
   
   This change adds a bounded range-based split that uses the measured `min` / 
`max` of the split key directly, falling back to a single full-table split when 
the range is unsafe to measure or would produce too many chunks.
   
   ## Changes
   
   - `DynamicChunkSplitter`:
     - When `queryApproximateRowCnt` returns `0` or a negative value, call the 
new `splitEvenlySizedChunksByRange` instead of the uneven-distribution path.
     - `splitEvenlySizedChunksByRange` generates chunks of `chunkSize` keys 
anchored at the measured `min`, stopping when the next boundary would exceed 
`max`, the chunk boundary stops advancing, or arithmetic overflows.
     - `isRangeChunkFallbackSafe` caps the estimated chunk count at the 
existing `split.sample-sharding.threshold` and refuses to fall back when `min`, 
`max` is an unsupported numeric pair (for example `NaN`).
   - `ObjectUtils#plus`: handle `Byte` and `Short` operands with 
`Math.addExact` and an explicit range check, so the new fallback works for all 
numeric split key types supported by `ObjectUtils#minus`.
   - `DynamicChunkSplitterTest`: use `ObjectUtils.compare` in the chunk 
ordering check, since the new test exercises `Byte` and `Short` ranges; add 
`testSplitEvenlySizedChunksByRangeWhenApproximateRowCountUnavailable` covering 
the happy path, oversized sparse ranges, unsafe numeric ranges, and the 
chunkSize / maxChunkCount argument validation.
   
   ## Validation
   
   ```
   ./mvnw -pl seatunnel-connectors-v2/connector-jdbc \
     -Dtest=DynamicChunkSplitterTest \
     -Dcheckstyle.skip -Dspotless.check.skip test
   ```
   
   Result: `Tests run: 6, Failures: 0, Errors: 0, Skipped: 0`
   
   ## Impact
   
   - Behavior change: when `approximateRowCnt <= 0` for a numeric split key, 
the source now uses a single range-based split plan rather than N boundary 
probes. Read SQL is unchanged; only the number and shape of splits differ. 
Tables with valid row count estimates are unaffected.
   - No new public API, no configuration change. The existing 
`split.sample-sharding.threshold` knob is reused as the chunk-count cap.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to