[ 
https://issues.apache.org/jira/browse/CAMEL-25512?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

shashank reassigned CAMEL-25512:
--------------------------------

    Assignee: shashank

> camel-minio - consumer routes an object again while its first exchange is 
> still in flight (asynchronous routes)
> ---------------------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25512
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25512
>             Project: Camel
>          Issue Type: Bug
>          Components: camel-minio
>            Reporter: shashank
>            Assignee: shashank
>            Priority: Major
>
> {{MinioConsumer.processBatch}} hands each exchange to the route with 
> {{getAsyncProcessor().process(exchange, EmptyAsyncCallback.get())}} (line 303 
> at main c578a42a776d) and returns without waiting; the object is deleted 
> ({{deleteAfterRead=true}}, the default) or moved only by the exchange's on 
> completion ({{processCommit}}, lines 288-301). When the route continues 
> asynchronously, for example {{to("seda:...")}} (seda hands the on completion 
> over to its consumer), an asynchronous producer such as kafka, or 
> {{threads()}}, the poll thread returns, and the next poll (500 ms later by 
> default) lists the same object again and routes it a second time. If the 
> first exchange deletes the object in between, the {{statObject}} in 
> {{createExchanges}} or the {{getObject}} in {{processBatch}} of the next poll 
> fails: that poll ends with an exception (logged as a WARN) and the objects it 
> had not routed yet wait for a later poll.
> This is the defect Claus described for aws2-s3 in CAMEL-17110 ("What the s3 
> consumer lacks is an in-progress repository"); that consumer has had one 
> since 4.7.0, the minio consumer does not. The same happens with 
> {{objectName}} set (one object per poll).
> camel-minio is deprecated in 4.23 (CAMEL-24750), but it is not deprecated in 
> the 4.14.x, 4.18.x and 4.22.x lines, which have the same code.
> h3. Reproduction
> New {{MinioConsumerInProgressTest}} with a minimal in-process fake of the S3 
> API calls the consumer makes (JDK {{HttpServer}}, no container; the MinIO ITs 
> are disabled since the images were removed): one object, route 
> {{minio:bkt?...}} to {{seda:work}}; the seda route holds the exchange until 
> the consumer has completed a second poll (counted with a {{pollStrategy}}). 
> Two cases: listing the bucket, and {{objectName=a.txt}}. On main both fail:
> {noformat}
> a.txt was consumed again by a later poll while its first exchange was still 
> in flight ==> expected: <1> but was: <2>
> {noformat}
> h3. Proposed fix
> The consumer keeps the names of the objects whose exchanges are in progress 
> (a set in the consumer): a poll skips them, in both listing and 
> {{objectName}} mode; a name is released when its exchange completes or fails 
> (in the on completion, after the delete/move), and when the poll does not 
> process it (stat/get failure, consumer stopping). No new option: aws2-s3 
> exposes a pluggable {{inProgressRepository}} endpoint option, which a 
> deprecated component should not gain; like aws2-s3's default memory 
> repository, the set is not cleared when the route is stopped, and an exchange 
> that never completes keeps its object reserved. With 
> {{deleteAfterRead=false}} an object is now consumed again only after its 
> previous exchange has completed (before, every poll routed it again, also 
> while it was still in flight). New {{MinioConsumerInProgressFailureTest}} 
> checks that a failed object is consumed again (passes on main too; it guards 
> the release on failure). camel-minio unit tests: 6, 0 failures.
> Not covered (as in aws2-s3): if the first exchange completes (deletes and 
> releases the object) after the next poll listed it but before that poll 
> reaches it, the stat of that poll still fails; this window is the time 
> between the list request and the stat of the object, instead of the whole 
> route.
> Found with a TLA+ model of the poll (list, stat per object, get + dispatch 
> per exchange) and asynchronous completion (delete): "with deleteAfterRead 
> every object is routed once" is violated in 10 steps (poll, dispatch, poll 
> again, dispatch again) and "a poll does not fail because of an object another 
> exchange is still processing" in 11 steps; both hold with a synchronous 
> route. A review model at statement granularity (the in-progress check per 
> object after the listing, delete before release) confirms that with the fix 
> no object is routed twice, and shows the remaining stat window above (12 
> steps). Then confirmed with the real consumer as above.
> Affected: main, camel-4.22.x, camel-4.18.x, camel-4.14.x (same code). The 
> google-storage consumer has the same shape (no in-progress set, delete in the 
> on completion, asynchronous hand-off; code reading only).
> Duplicate check (2026-10-09): JIRA component camel-minio (18 issues; 
> CAMEL-17100 is about the listing speed), text "minio" since 2024 (11), 
> "in-progress repository" / "maxMessagesPerPoll" since 2023 (26; CAMEL-24491 
> is the ibm-cos sibling); GitHub pull requests "minio", "minio consumer": none 
> open about the consumer; no open PR touches camel-minio.
> _Filed with Claude Code on behalf of allthingssecurity._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to