Andrea Cosentino created CAMEL-25169:
----------------------------------------
Summary: camel-couchbase: the Couchbase connection is never closed
- core().shutdown() returns a cold Mono that is discarded
Key: CAMEL-25169
URL: https://issues.apache.org/jira/browse/CAMEL-25169
Project: Camel
Issue Type: Bug
Components: camel-couchbase
Reporter: Andrea Cosentino
Assignee: Andrea Cosentino
Neither the consumer nor the producer ever closes the Couchbase connection it
opened, so every route stop/start leaks a {{Cluster}} and a
{{ClusterEnvironment}} - with their Netty event loops, schedulers and sockets -
for the lifetime of the JVM.
h2. What the code does
{{CouchbaseConsumer.doStop()}} and {{CouchbaseProducer.doShutdown()}} both do
the same thing:
{code:java}
if (client != null) {
client.core().shutdown();
}
{code}
{{com.couchbase.client.core.Core.shutdown()}} returns
{{reactor.core.publisher.Mono<Void>}}. Disassembling {{core-io-3.12.3}},
{{shutdown(Duration)}} is assembled purely from {{Mono.fromRunnable(...)}},
{{.then(...)}} and {{Mono.defer(...)}} - it is a *cold* publisher. The returned
{{Mono}} is discarded, nothing subscribes to it, and therefore nothing is shut
down. The call compiles, reads like a close, and has no effect at all.
h2. Why the leak is per-route and not per-component
{{CouchbaseEndpoint.createClient()}} builds a new {{ClusterEnvironment}} *and*
a new {{Cluster}} on every call, and is called once from {{createProducer()}}
and once from {{createConsumer(...)}}:
{code:java}
Cluster cluster = Cluster.connect(connStr, ClusterOptions
.clusterOptions(username, password)
.environment(env));
return cluster.bucket(bucket);
{code}
The {{Cluster}} handle is dropped on the floor - only the {{Bucket}} is
returned - so even a caller that wanted to close it has nothing to close.
There is a second half to this. Because Camel supplies its own environment
through {{ClusterOptions.environment(env)}},
{{AsyncCluster.extractClusterEnvironment}} registers it as
{{OwnedOrExternal.external(...)}} rather than {{owned}}. An externally supplied
environment is *not* shut down by {{Cluster.disconnect()}} - that is the SDK's
documented contract, and it is visible in the bytecode. So the fix has to
disconnect the cluster *and* call {{ClusterEnvironment.shutdown()}}; doing only
the first would still leak the event loops.
A route using {{toD("couchbase:...")}} over varying URIs leaks one pair per
distinct URI.
h2. History - this was fixed once already
CAMEL-11674 (2017), titled "Couchbase client is never shut down", fixed exactly
this by adding {{client.shutdown()}} to both the consumer and the producer. On
the 2.x SDK that was a blocking call and it worked.
The 3.0.5 SDK upgrade (CAMEL-14319, commit {{025549c010e4}}, 2020-06-26)
mechanically rewrote it to {{bucket.core().shutdown()}} as part of a large API
migration. The migration was not wrong to look for a replacement - it just
landed on one that is lazy. The 2017 fix has been inert ever since.
h2. Proposed fix
Have {{CouchbaseEndpoint}} retain the {{Cluster}} and the
{{ClusterEnvironment}} it creates and expose a close, then call it from
{{CouchbaseConsumer.doStop()}} and {{CouchbaseProducer.doShutdown()}}:
{{cluster.disconnect()}} followed by {{environment.shutdown()}}.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)