dybyte commented on PR #11960:
URL: https://github.com/apache/seatunnel/pull/11960#issuecomment-5410684809

   > > I think this overlaps with some of the work already being done in 
#11912, especially around replacing `LocalSchemaCoordinator` and 
checkpoint-based schema coordination. Could you clarify how you see the 
relationship between this issue and #11912? Is this intended as an alternative 
design, or as additional work on top of it?
   > 
   > It seems you focused on rewriting the SinkWriter layer. I replaced the 
entire Flink Coordinator, primarily addressing fault recovery issues by 
replacing the coordinator with a data plane protocol. I consider this 
additional work, and you can focus on the SinkWriter side. Your implementation 
essentially still utilizes the coordinator, while I completely removed it, 
using Flink's native mechanisms for coordination. In this branch, I didn't make 
any additional rewrites to the SinkWriter.
   
   Thanks for the clarification. Just to clarify one point about #11912: the 
current implementation is not limited to the SinkWriter layer, and it no longer 
uses LocalSchemaCoordinator. SchemaOperator also removes requestSchemaChange() 
and relies on Flink checkpoint completion to coordinate schema-change dispatch 
and the release of buffered rows.
   
   I agree that #11960 adds stronger recovery/rescaling mechanisms, such as 
explicit producer/sequence IDs, atomic protocol state, and replay handling.
   
   However, I think there is still an architectural overlap between the two 
PRs. #11912 keeps parallel sink writers for the same table and separates 
one-time external DDL application from writer-local schema refresh, while 
#11960 appears to partition data by table and let one table owner handle the 
schema change. So these seem less like independent SinkWriter/coordinator 
changes and more like two different coordination models for part of the same 
problem.
   
   Would it make sense to first agree on which parallel-writer model we want to 
keep, and then treat the additional recovery/rescaling guarantees in #11960 as 
follow-up hardening on top of that model?


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to