FrankChen021 commented on PR #20382: URL: https://github.com/apache/druid/pull/20382#issuecomment-5770275619
> > Finding: When druid.storage.zip=false, pushNoZip uploads each file directly under the final segment prefix and only then lists and deletes stale objects. If a historical or another pusher reads the same path during a retry or same-path rewrite, it can observe missing new files, old chunks, or a mixture of two segment versions; the puller can then materialize a corrupt or incomplete segment even though the push later succeeds. > > I don't think this is really a problem, is it? Historicals won't ever be pulling a segment that didn't publish successfully (load spec doesn't exist in metadata until pushes all suceed), ingest time replicas I don't believe will publish at the same time (and they should be producing identical contents), and the only real replace case I'm aware of where content could be different is like when streaming failed to publish and has to retry and while starting at the same offsets, might end at different offsets, but that runs with useUniquePath=true so avoid the problem. You're right. Sometimes AI does not correctly under the business flow. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
