Freeman Yue Fang created CXF-9255:
-------------------------------------
Summary: WS-RM: retransmissions of messages cached in the same
millisecond may run out of order and stall in-order delivery until the receive
timeout
Key: CXF-9255
URL: https://issues.apache.org/jira/browse/CXF-9255
Project: CXF
Issue Type: Improvement
Components: WS-* Components
Reporter: Freeman Yue Fang
{panel:title=Symptom}
The in-order one-way tests in {{systests/ws-rm}} time out sometimes, and more
frequently on a fast machine :
- {{DeliveryAssuranceOnewayTest.testExactlyOnceInOrder}},
{{testExactlyOnceInOrderAsyncExecutor}}, {{testInOrder}},
{{testInOrderAsyncExecutor}}
- {{MessageCallbackOnewayTest.testExactlyOnceInOrder}}
{noformat}
DeliveryAssuranceOnewayTest.testExactlyOnceInOrder:327->testOnewayExactlyOnceInOrder:343
ยป TestTimedOut test timed out after 60000 milliseconds
{noformat}
Each test passes when run on its own. Whether it fails depends on timing, not
on test order.
{panel}
h3. Root cause
{{RetransmissionQueueImpl.ResendCandidate}} schedules its first resend at
{{System.currentTimeMillis() + baseRetransmissionInterval}}. Each candidate
gets its own {{TimerTask}} on the single shared {{java.util.Timer}} of the
{{RMManager}}. {{java.util.Timer}} gives no ordering guarantee for tasks due at
the same time.
When two messages of the same sequence are cached in the same millisecond,
which is common on fast hardware, the resend of the later message can run
first. Using the tests' scenario (message 2 lost by {{MessageLossSimulator}},
in-order delivery):
Message 3 is held by the destination until message 2 arrives, so the client
call for message 3 waits for its HTTP response.
The timer fires the resend of message 3 before the resend of message 2. The
resend runs synchronously on the timer thread ({{SynchronousExecutor}}), or on
the client's single-thread executor in the {{*AsyncExecutor}} variants, and
blocks for the same reason.
Message 2 is never resent while that thread is blocked. Everything waits for
the HTTP receive timeout (60 s by default) before message 2 is finally resent.
In production this means the same thing can happen outside the tests: with
in-order delivery, a later message resent ahead of an earlier one stalls
retransmission until the receive timeout.
h3. Evidence
Thread dumps of a hung test show the client blocked in
{{HttpClientHTTPConduit.getResponse()}} while all of the server's Jetty threads
are idle. RM FINE logs show the resend order in every run:
||Run||Messages 2 and 3 cached||Resend order||Result||
|failed (4 of 4 runs)|same millisecond (e.g. .641 / .641)|1, 3 (2 only ~60 s
later)|timeout|
|passed|different milliseconds (e.g. .396 / .397)|1, 2, 3|~4 s|
Switching to the {{URLConnection}} conduit, or turning off HttpClient sharing,
doesn't change the result.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)