Andrea Cosentino created CAMEL-25169:
----------------------------------------

             Summary: camel-couchbase: the Couchbase connection is never closed 
- core().shutdown() returns a cold Mono that is discarded
                 Key: CAMEL-25169
                 URL: https://issues.apache.org/jira/browse/CAMEL-25169
             Project: Camel
          Issue Type: Bug
          Components: camel-couchbase
            Reporter: Andrea Cosentino
            Assignee: Andrea Cosentino


Neither the consumer nor the producer ever closes the Couchbase connection it 
opened, so every route stop/start leaks a {{Cluster}} and a 
{{ClusterEnvironment}} - with their Netty event loops, schedulers and sockets - 
for the lifetime of the JVM.

h2. What the code does

{{CouchbaseConsumer.doStop()}} and {{CouchbaseProducer.doShutdown()}} both do 
the same thing:

{code:java}
if (client != null) {
    client.core().shutdown();
}
{code}

{{com.couchbase.client.core.Core.shutdown()}} returns 
{{reactor.core.publisher.Mono<Void>}}. Disassembling {{core-io-3.12.3}}, 
{{shutdown(Duration)}} is assembled purely from {{Mono.fromRunnable(...)}}, 
{{.then(...)}} and {{Mono.defer(...)}} - it is a *cold* publisher. The returned 
{{Mono}} is discarded, nothing subscribes to it, and therefore nothing is shut 
down. The call compiles, reads like a close, and has no effect at all.

h2. Why the leak is per-route and not per-component

{{CouchbaseEndpoint.createClient()}} builds a new {{ClusterEnvironment}} *and* 
a new {{Cluster}} on every call, and is called once from {{createProducer()}} 
and once from {{createConsumer(...)}}:

{code:java}
Cluster cluster = Cluster.connect(connStr, ClusterOptions
        .clusterOptions(username, password)
        .environment(env));

return cluster.bucket(bucket);
{code}

The {{Cluster}} handle is dropped on the floor - only the {{Bucket}} is 
returned - so even a caller that wanted to close it has nothing to close.

There is a second half to this. Because Camel supplies its own environment 
through {{ClusterOptions.environment(env)}}, 
{{AsyncCluster.extractClusterEnvironment}} registers it as 
{{OwnedOrExternal.external(...)}} rather than {{owned}}. An externally supplied 
environment is *not* shut down by {{Cluster.disconnect()}} - that is the SDK's 
documented contract, and it is visible in the bytecode. So the fix has to 
disconnect the cluster *and* call {{ClusterEnvironment.shutdown()}}; doing only 
the first would still leak the event loops.

A route using {{toD("couchbase:...")}} over varying URIs leaks one pair per 
distinct URI.

h2. History - this was fixed once already

CAMEL-11674 (2017), titled "Couchbase client is never shut down", fixed exactly 
this by adding {{client.shutdown()}} to both the consumer and the producer. On 
the 2.x SDK that was a blocking call and it worked.

The 3.0.5 SDK upgrade (CAMEL-14319, commit {{025549c010e4}}, 2020-06-26) 
mechanically rewrote it to {{bucket.core().shutdown()}} as part of a large API 
migration. The migration was not wrong to look for a replacement - it just 
landed on one that is lazy. The 2017 fix has been inert ever since.

h2. Proposed fix

Have {{CouchbaseEndpoint}} retain the {{Cluster}} and the 
{{ClusterEnvironment}} it creates and expose a close, then call it from 
{{CouchbaseConsumer.doStop()}} and {{CouchbaseProducer.doShutdown()}}: 
{{cluster.disconnect()}} followed by {{environment.shutdown()}}.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to