shashank created CAMEL-25505:
--------------------------------

             Summary: camel-lucene - inserts fail for good after the route is 
restarted, after one failed insert, and when two inserts run at the same time
                 Key: CAMEL-25505
                 URL: https://issues.apache.org/jira/browse/CAMEL-25505
             Project: Camel
          Issue Type: Bug
          Components: camel-lucene
            Reporter: shashank


A {{lucene:<name>:insert}} endpoint creates one {{LuceneIndexer}} (one 
{{NIOFSDirectory}} for the index) in its constructor ({{LuceneEndpoint}} line 
53 at main c578a42a776d), and every producer of the endpoint shares it. Three 
ways this breaks indexing:

* {{LuceneIndexProducer.doStop()}} closes the shared directory (line 35). After 
the route is stopped and started again (route controller, JMX, supervising 
controller, route reload), or after any other route that sends to the same 
endpoint is stopped, every insert fails with 
{{org.apache.lucene.store.AlreadyClosedException: this Directory is closed}} 
until the CamelContext is restarted.
* {{LuceneIndexer.index()}} opens an {{IndexWriter}} (which takes the index 
{{write.lock}}), adds the documents and only then commits and closes it (lines 
75-85). An exchange that fails in between, for example one without a body 
({{getMandatoryBody}}) or with a header that cannot be converted to a String, 
leaves the writer open: the write lock stays held in the JVM and every later 
insert fails with {{LockObtainFailedException: Lock held by this virtual 
machine}}, until the JVM is restarted.
* {{index()}} is not synchronized and keeps the writer in a field, so two 
exchanges indexing at the same time (two routes, or a route with concurrent 
consumers) fail on the same lock.

h3. Reproduction

New {{LuceneIndexProducerLifecycleTest}} (two routes {{direct:index}} and 
{{direct:other}} to the same insert endpoint on a {{@TempDir}} index; no 
sleeps, latches and Awaitility). On main all four fail:
{noformat}
insertWorksAfterRouteRestart: an insert after the route was restarted must 
succeed ==> expected: <null> but was: 
<org.apache.lucene.store.AlreadyClosedException: this Directory is closed>
stoppingOneRouteDoesNotBreakAnotherRouteOnTheSameEndpoint: ... expected: <null> 
but was: <org.apache.lucene.store.AlreadyClosedException: this Directory is 
closed>
failedInsertDoesNotBlockLaterInserts: a failed insert must not leave the index 
write lock held ==> expected: <null> but was: 
<org.apache.lucene.store.LockObtainFailedException: Lock held by this virtual 
machine: .../write.lock>
concurrentInsertsAreIndexed: (a header whose String conversion waits while the 
first insert holds the writer) ... but was: 
<org.apache.lucene.store.LockObtainFailedException: Lock held by this virtual 
machine: .../write.lock>
{noformat}

h3. Proposed fix

* The endpoint closes the directory in {{doShutdown()}}; the producer no longer 
closes it in {{doStop()}}.
* {{LuceneIndexer.index()}} is {{synchronized}} and calls 
{{indexWriter.rollback()}} (discards the documents of the failed exchange, 
closes the writer, releases the lock) when adding fails, then rethrows.

No behaviour change for a working route except that concurrent inserts to one 
endpoint now wait for each other instead of failing. camel-lucene tests with 
the fix: {{LuceneIndexProducerLifecycleTest}} 4, 
{{LuceneIndexAndQueryProducerIT}} 4, {{LuceneQueryProcessorIT}} 2, 0 failures.

Found with a TLA+ model of the insert path (producer start/stop per route, the 
shared directory, the write lock and the indexWriter field, one action per Java 
statement that touches them): "every exchange sent to a started insert producer 
is indexed" is violated in 4 steps (stop, start, send, open fails on the closed 
directory), in 3 steps with two routes, in 4 steps with two concurrent senders, 
and, once the model got the failing-exchange path that the realizability check 
showed, in 4 steps after a failed exchange. With the fix the property holds for 
2 routes x 2 threads with 2 stops. Then confirmed with the real component as 
above.

Affected: main, camel-4.22.x, camel-4.18.x, camel-4.14.x (same code, GitHub 
contents API; unchanged since the component was added in 2.2.0).

Duplicate check (2026-10-09): JIRA component camel-lucene (8 issues, all 
upgrades or naming), text "LuceneSearcher" (4, upgrades), 
"AlreadyClosedException" (4, all camel-rabbitmq), "LockObtainFailedException" 
(none); GitHub pull requests "camel-lucene": none about the producers.

_Filed with Claude Code on behalf of allthingssecurity._




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to