shashank created CAMEL-25505:
--------------------------------
Summary: camel-lucene - inserts fail for good after the route is
restarted, after one failed insert, and when two inserts run at the same time
Key: CAMEL-25505
URL: https://issues.apache.org/jira/browse/CAMEL-25505
Project: Camel
Issue Type: Bug
Components: camel-lucene
Reporter: shashank
A {{lucene:<name>:insert}} endpoint creates one {{LuceneIndexer}} (one
{{NIOFSDirectory}} for the index) in its constructor ({{LuceneEndpoint}} line
53 at main c578a42a776d), and every producer of the endpoint shares it. Three
ways this breaks indexing:
* {{LuceneIndexProducer.doStop()}} closes the shared directory (line 35). After
the route is stopped and started again (route controller, JMX, supervising
controller, route reload), or after any other route that sends to the same
endpoint is stopped, every insert fails with
{{org.apache.lucene.store.AlreadyClosedException: this Directory is closed}}
until the CamelContext is restarted.
* {{LuceneIndexer.index()}} opens an {{IndexWriter}} (which takes the index
{{write.lock}}), adds the documents and only then commits and closes it (lines
75-85). An exchange that fails in between, for example one without a body
({{getMandatoryBody}}) or with a header that cannot be converted to a String,
leaves the writer open: the write lock stays held in the JVM and every later
insert fails with {{LockObtainFailedException: Lock held by this virtual
machine}}, until the JVM is restarted.
* {{index()}} is not synchronized and keeps the writer in a field, so two
exchanges indexing at the same time (two routes, or a route with concurrent
consumers) fail on the same lock.
h3. Reproduction
New {{LuceneIndexProducerLifecycleTest}} (two routes {{direct:index}} and
{{direct:other}} to the same insert endpoint on a {{@TempDir}} index; no
sleeps, latches and Awaitility). On main all four fail:
{noformat}
insertWorksAfterRouteRestart: an insert after the route was restarted must
succeed ==> expected: <null> but was:
<org.apache.lucene.store.AlreadyClosedException: this Directory is closed>
stoppingOneRouteDoesNotBreakAnotherRouteOnTheSameEndpoint: ... expected: <null>
but was: <org.apache.lucene.store.AlreadyClosedException: this Directory is
closed>
failedInsertDoesNotBlockLaterInserts: a failed insert must not leave the index
write lock held ==> expected: <null> but was:
<org.apache.lucene.store.LockObtainFailedException: Lock held by this virtual
machine: .../write.lock>
concurrentInsertsAreIndexed: (a header whose String conversion waits while the
first insert holds the writer) ... but was:
<org.apache.lucene.store.LockObtainFailedException: Lock held by this virtual
machine: .../write.lock>
{noformat}
h3. Proposed fix
* The endpoint closes the directory in {{doShutdown()}}; the producer no longer
closes it in {{doStop()}}.
* {{LuceneIndexer.index()}} is {{synchronized}} and calls
{{indexWriter.rollback()}} (discards the documents of the failed exchange,
closes the writer, releases the lock) when adding fails, then rethrows.
No behaviour change for a working route except that concurrent inserts to one
endpoint now wait for each other instead of failing. camel-lucene tests with
the fix: {{LuceneIndexProducerLifecycleTest}} 4,
{{LuceneIndexAndQueryProducerIT}} 4, {{LuceneQueryProcessorIT}} 2, 0 failures.
Found with a TLA+ model of the insert path (producer start/stop per route, the
shared directory, the write lock and the indexWriter field, one action per Java
statement that touches them): "every exchange sent to a started insert producer
is indexed" is violated in 4 steps (stop, start, send, open fails on the closed
directory), in 3 steps with two routes, in 4 steps with two concurrent senders,
and, once the model got the failing-exchange path that the realizability check
showed, in 4 steps after a failed exchange. With the fix the property holds for
2 routes x 2 threads with 2 stops. Then confirmed with the real component as
above.
Affected: main, camel-4.22.x, camel-4.18.x, camel-4.14.x (same code, GitHub
contents API; unchanged since the component was added in 2.2.0).
Duplicate check (2026-10-09): JIRA component camel-lucene (8 issues, all
upgrades or naming), text "LuceneSearcher" (4, upgrades),
"AlreadyClosedException" (4, all camel-rabbitmq), "LockObtainFailedException"
(none); GitHub pull requests "camel-lucene": none about the producers.
_Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)