zhangshenghang opened a new pull request, #12525:
URL: https://github.com/apache/seatunnel/pull/12525

   ### Purpose of this pull request
   
   Close #12521.
   
   `AbstractSchema#indexOf`, `#getColumn` and `#contains` scan the 
`columnNames` list linearly on every call. These lookups sit on hot paths (for 
example sink writers resolving field indexes per table), and for wide tables 
with hundreds of columns the repeated O(n) scans add significant overhead.
   
   This PR builds a lazily-initialized name-to-index map (double-checked, 
thread safe) and uses it for the lookups:
   
   - `indexOf`/`contains` become hash lookups; `getColumn` reuses `indexOf` and 
keeps its previous semantics (including the exception behavior for missing 
columns).
   - Duplicate column names keep the first occurrence, matching the former 
linear-scan semantics.
   - The constructor now takes a defensive copy of the column list 
(unmodifiable), so external mutation of the caller's list can no longer change 
the schema or invalidate the cache.
   
   ### Does this PR introduce _any_ user-facing change?
   
   No behavior change is intended. The only observable difference is that 
mutating the list passed to the schema constructor after construction no longer 
affects the schema (previously an undocumented aliasing hazard), and 
`getColumns()` now returns an unmodifiable view of an internal copy.
   
   ### How was this patch tested?
   
   - Added `AbstractSchemaTest` covering lookups, duplicate-name first-match 
semantics, and the defensive copy.
   - `mvn -pl seatunnel-api test`: 404 tests, all passing.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to