Justin, thanks for replying. Now I’ve got the point of such behavior. Yes, that is exactly what I meant. If you expect any connection loss between HA pair while having client able to connect to both, you may encounter message loss, and replication is not a right solution.
> 27 мая 2022 г., в 01:42, Justin Bertram <[email protected]> написал(а): > >> But in case of replication failure, for example network failure it will > not fail a transaction. > > I'm not exactly sure what you mean here. Are you saying that if a primary > is replicating to a backup and a client is in the middle of a transaction > and the replication connection between the primary and backup fails then > the client's transaction will complete successfully? If so, that's the > behavior I would expect. The loss of the replication connection potentially > represents the crashing of the backup. At the very least it represents some > kind of network problem. In any case, there is nothing wrong with the > primary broker in this circumstance so there is no reason to fail the > client's transaction. If failures with the backup caused the primary broker > to fail client operations that would otherwise succeed then adding a backup > would *increase* the likelihood of client-facing failures instead of > decreasing them. This is typically the opposite of what is wanted when > configuring HA. > > I can imagine use-cases where it is absolutely critical for data to be > replicated successfully and any failure to replicate should be considered > fatal for any client operation. However, that behavior is not supported via > replication. You would need to use shared-storage to get this kind of > behavior. > > > Justin > > On Thu, May 26, 2022 at 4:51 PM Илья Грушевский <[email protected]> wrote: > >> You are right, I should have not use the term asynchronous. >> But in case of replication failure, for example network failure it will >> not fail a transaction. >> So if I gradually lose connections, first between primary and backup and >> then between primary and client I will lose all send message between those >> events. >> In case of cluster I may lose connection between primary and client if >> primary node decides to turn itself off after quorum vote. >> >>> 27 мая 2022 г., в 00:31, Justin Bertram <[email protected]> >> написал(а): >>> >>>> I think this is due to the fact that HA replication is asynchronous and >>> replica server may not catch up with primary. >>> >>> To be clear, message replication between a primary and a backup is >>> *synchronous*. >>> >>> >>> Justin >>> >>> On Thu, May 26, 2022 at 3:11 AM Iliya Grushevskiy <[email protected]> >>> wrote: >>> >>>> Hi, Aaron >>>> >>>> We are currently testing similar deployment and have encountered several >>>> issues: >>>> >>>> - message lose on send on network failure between data centers >>>> I think this is due to the fact that HA replication is asynchronous and >>>> replica server may not catch up with primary. >>>> >>>> - message lose or duplicate (depending on error handling strategy) on >>>> consumer on network failure between data centers >>>> I think this was caused by two factors: duplicate id cache is >> consistent >>>> only in HA pair and message redistribution was on. >>>> Switching off redistribution (or as an option increasing delay) should >>>> fix this issue. >>>> >>>> - message duplicate on mirrored server >>>> This is addressed in pull request: >>>> https://github.com/apache/activemq-artemis/pull/4066 >>>> >>>> Regards >>>> Iliya Grushevskiy >>>> >>>> >>>>> 26 мая 2022 г., в 07:46, Justin Bertram <[email protected]> >>>> написал(а): >>>>> >>>>> I'm not aware of such a production deployment and I would be surprised >> if >>>>> there was one given that clustering was designed for local area >> networks >>>>> with low latency which typically isn't what is found between data >>>> centers. >>>>> >>>>> I recommend you pursue your mirroring approach as that is what >> mirroring >>>>> was designed for (i.e. cross data-center disaster-recovery use-cases). >>>>> >>>>> >>>>> Justin >>>>> >>>>> On Wed, May 25, 2022 at 10:36 PM Steigerwald, Aaron >>>>> <[email protected]> wrote: >>>>> >>>>>> Hello, >>>>>> >>>>>> Is anyone aware of a production deployment of an Artemis "cross data >>>>>> center" HA cluster? For example, a cluster spread across 3 data >> centers. >>>>>> Each data center contains a master/slave pair. >>>>>> >>>>>> I would like to know what kind of issues anyone has overcome with >> such a >>>>>> configuration. I understand there are many configuration and >> operational >>>>>> variables. Any info would be helpful. >>>>>> >>>>>> Note that we are considering asynchronously mirroring each >> master/slave >>>>>> pair's queues to a dedicated asynchronous target node. The >> asynchronous >>>>>> target node would exist in a different data center and would not >> service >>>>>> any other connections. A custom plugin would automatically scale down >>>> the >>>>>> messages into a live cluster node if the connections to the >> master/slave >>>>>> mirror sources were disconnected for a period of time. >>>>>> >>>>>> Thank you, >>>>>> Aaron Steigerwald >>>>>> >>>> >>>> >> >>
