[ 
https://issues.apache.org/jira/browse/CXF-9255?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Freeman Yue Fang resolved CXF-9255.
-----------------------------------
    Fix Version/s: 4.2.5
                   4.1.10
       Resolution: Fixed

> WS-RM: retransmissions of messages cached in the same millisecond may run out 
> of order and stall in-order delivery until the receive timeout
> --------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: CXF-9255
>                 URL: https://issues.apache.org/jira/browse/CXF-9255
>             Project: CXF
>          Issue Type: Improvement
>          Components: WS-* Components
>            Reporter: Freeman Yue Fang
>            Assignee: Freeman Yue Fang
>            Priority: Major
>             Fix For: 4.2.5, 4.1.10
>
>
> {panel:title=Symptom}
> The in-order one-way tests in {{systests/ws-rm}} time out sometimes, and more 
> frequently on a fast machine :
> - {{DeliveryAssuranceOnewayTest.testExactlyOnceInOrder}}, 
> {{testExactlyOnceInOrderAsyncExecutor}}, {{testInOrder}}, 
> {{testInOrderAsyncExecutor}}
> - {{MessageCallbackOnewayTest.testExactlyOnceInOrder}}
> {noformat}
> DeliveryAssuranceOnewayTest.testExactlyOnceInOrder:327->testOnewayExactlyOnceInOrder:343
>  ยป TestTimedOut test timed out after 60000 milliseconds
> {noformat}
> Each test passes when run on its own. Whether it fails depends on timing, not 
> on test order.
> {panel}
> h3. Root cause
> {{RetransmissionQueueImpl.ResendCandidate}} schedules its first resend at 
> {{System.currentTimeMillis() + baseRetransmissionInterval}}. Each candidate 
> gets its own {{TimerTask}} on the single shared {{java.util.Timer}} of the 
> {{RMManager}}. {{java.util.Timer}} gives no ordering guarantee for tasks due 
> at the same time.
> When two messages of the same sequence are cached in the same millisecond, 
> which is common on fast hardware, the resend of the later message can run 
> first. Using the tests' scenario (message 2 lost by {{MessageLossSimulator}}, 
> in-order delivery):
> Message 3 is held by the destination until message 2 arrives, so the client 
> call for message 3 waits for its HTTP response.
> The timer fires the resend of message 3 before the resend of message 2. The 
> resend runs synchronously on the timer thread ({{SynchronousExecutor}}), or 
> on the client's single-thread executor in the {{*AsyncExecutor}} variants, 
> and blocks for the same reason.
> Message 2 is never resent while that thread is blocked. Everything waits for 
> the HTTP receive timeout (60 s by default) before message 2 is finally resent.
> In production this means the same thing can happen outside the tests: with 
> in-order delivery, a later message resent ahead of an earlier one stalls 
> retransmission until the receive timeout.
> h3. Evidence
> Thread dumps of a hung test show the client blocked in 
> {{HttpClientHTTPConduit.getResponse()}} while all of the server's Jetty 
> threads are idle. RM FINE logs show the resend order in every run:
> ||Run||Messages 2 and 3 cached||Resend order||Result||
> |failed (4 of 4 runs)|same millisecond (e.g. .641 / .641)|1, 3 (2 only ~60 s 
> later)|timeout|
> |passed|different milliseconds (e.g. .396 / .397)|1, 2, 3|~4 s|
> Switching to the {{URLConnection}} conduit, or turning off HttpClient 
> sharing, doesn't change the result.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to