[
https://issues.apache.org/jira/browse/CAMEL-25505?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Work on CAMEL-25505 started by shashank.
----------------------------------------
> camel-lucene - inserts fail for good after the route is restarted, after one
> failed insert, and when two inserts run at the same time
> -------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25505
> URL: https://issues.apache.org/jira/browse/CAMEL-25505
> Project: Camel
> Issue Type: Bug
> Components: camel-lucene
> Reporter: shashank
> Assignee: shashank
> Priority: Major
>
> A {{lucene:<name>:insert}} endpoint creates one {{LuceneIndexer}} (one
> {{NIOFSDirectory}} for the index) in its constructor ({{LuceneEndpoint}} line
> 53 at main c578a42a776d), and every producer of the endpoint shares it. Three
> ways this breaks indexing:
> * {{LuceneIndexProducer.doStop()}} closes the shared directory (line 35).
> After the route is stopped and started again (route controller, JMX,
> supervising controller, route reload), or after any other route that sends to
> the same endpoint is stopped, every insert fails with
> {{org.apache.lucene.store.AlreadyClosedException: this Directory is closed}}
> until the CamelContext is restarted.
> * {{LuceneIndexer.index()}} opens an {{IndexWriter}} (which takes the index
> {{write.lock}}), adds the documents and only then commits and closes it
> (lines 75-85). An exchange that fails in between, for example one without a
> body ({{getMandatoryBody}}) or with a header that cannot be converted to a
> String, leaves the writer open: the write lock stays held in the JVM and
> every later insert fails with {{LockObtainFailedException: Lock held by this
> virtual machine}}, until the JVM is restarted.
> * {{index()}} is not synchronized and keeps the writer in a field, so two
> exchanges indexing at the same time (two routes, or a route with concurrent
> consumers) fail on the same lock.
> h3. Reproduction
> New {{LuceneIndexProducerLifecycleTest}} (two routes {{direct:index}} and
> {{direct:other}} to the same insert endpoint on a {{@TempDir}} index; no
> sleeps, latches and Awaitility). On main all four fail:
> {noformat}
> insertWorksAfterRouteRestart: an insert after the route was restarted must
> succeed ==> expected: <null> but was:
> <org.apache.lucene.store.AlreadyClosedException: this Directory is closed>
> stoppingOneRouteDoesNotBreakAnotherRouteOnTheSameEndpoint: ... expected:
> <null> but was: <org.apache.lucene.store.AlreadyClosedException: this
> Directory is closed>
> failedInsertDoesNotBlockLaterInserts: a failed insert must not leave the
> index write lock held ==> expected: <null> but was:
> <org.apache.lucene.store.LockObtainFailedException: Lock held by this virtual
> machine: .../write.lock>
> concurrentInsertsAreIndexed: (a header whose String conversion waits while
> the first insert holds the writer) ... but was:
> <org.apache.lucene.store.LockObtainFailedException: Lock held by this virtual
> machine: .../write.lock>
> {noformat}
> h3. Proposed fix
> * The endpoint closes the directory in {{doShutdown()}}; the producer no
> longer closes it in {{doStop()}}.
> * {{LuceneIndexer.index()}} is {{synchronized}} and calls
> {{indexWriter.rollback()}} (discards the documents of the failed exchange,
> closes the writer, releases the lock) when adding fails, then rethrows.
> No behaviour change for a working route except that concurrent inserts to one
> endpoint now wait for each other instead of failing. camel-lucene tests with
> the fix: {{LuceneIndexProducerLifecycleTest}} 4,
> {{LuceneIndexAndQueryProducerIT}} 4, {{LuceneQueryProcessorIT}} 2, 0 failures.
> Found with a TLA+ model of the insert path (producer start/stop per route,
> the shared directory, the write lock and the indexWriter field, one action
> per Java statement that touches them): "every exchange sent to a started
> insert producer is indexed" is violated in 4 steps (stop, start, send, open
> fails on the closed directory), in 3 steps with two routes, in 4 steps with
> two concurrent senders, and, once the model got the failing-exchange path
> that the realizability check showed, in 4 steps after a failed exchange. With
> the fix the property holds for 2 routes x 2 threads with 2 stops. Then
> confirmed with the real component as above.
> Affected: main, camel-4.22.x, camel-4.18.x, camel-4.14.x (same code, GitHub
> contents API; unchanged since the component was added in 2.2.0).
> Duplicate check (2026-10-09): JIRA component camel-lucene (8 issues, all
> upgrades or naming), text "LuceneSearcher" (4, upgrades),
> "AlreadyClosedException" (4, all camel-rabbitmq), "LockObtainFailedException"
> (none); GitHub pull requests "camel-lucene": none about the producers.
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)