Croway opened a new pull request, #27132:
URL: https://github.com/apache/camel/pull/27132

   [CAMEL-25167](https://issues.apache.org/jira/browse/CAMEL-25167)
   
   ## Problem
   
   Object-storage producers each have their own way to find the length of a 
body that is not a `java.io.File`, and most of them either fail or load the 
whole body into memory:
   
   - **aws2-s3 and azure-storage-blob** probe the length with `mark(1024)` / 
`skip(available())` / `reset()`. On a `BufferedInputStream` past the mark 
buffer (about 8 KB) this throws `IOException: Resetting to invalid mark`. A 
`java.nio.file.Path` body is converted to exactly such a stream 
(`IOConverter.toInputStream(Path)`). For example, a Spring Boot platform-http 
multipart upload (the body is a `Path`) sent to `aws2-s3` fails for anything 
over a few KB.
   - **azure-storage-datalake** rejects streams without mark support. For the 
others it reads the whole stream into a `byte[]` to measure it.
   - **aws2-s3, minio, google-storage and ibm-cos** read a body of unknown 
length fully into a `ByteArrayOutputStream`. That bypasses stream caching, so 
spooling to disk never applies.
   
   The ftp, sftp and smb producers with `charset` set do 
`getMandatoryBody(String.class).getBytes(charset)`, which holds the whole file 
in memory twice.
   
   ## Changes
   
   - **camel-support:** new `PayloadHelper`.
     - `getBodyLength(Message)` / `getLength(Object)` return the length without 
reading the body: `WrappedFile` length, `File` / `Path` size, 
`StreamCache.length()`, `byte[]`, `ByteBuffer`, `ByteArrayInputStream`, 
`FileInputStream`. Otherwise they return `-1`. There is no mark/skip/reset 
probing.
     - `cacheStream(Exchange, InputStream)` copies a stream of unknown length 
through `OutputStreamBuilder`, so it is spooled to disk when stream-caching 
spooling is enabled. With spooling off, it behaves as before (in memory).
   - **aws2-s3, azure-storage-blob, azure-storage-datalake, minio, ibm-cos, 
google-storage:** use `PayloadHelper` instead of the per-component probes and 
in-memory copies.
     - A `Path` body now has a known length, so it is uploaded directly.
     - azure-storage-datalake accepts streams without mark support.
     - azure blob and datalake now convert a remote wrapped file (e.g. from 
SFTP) to a stream; before, they tried to convert the remote file handle.
     - `AWS2S3Utils.determineLengthInputStream`, 
`BlobUtils.getInputStreamLength` and `DataLakeUtils.getInputStreamLength` stay 
public and delegate to `PayloadHelper`.
   - **camel-util:** new `IOHelper.ReaderInputStream`, which encodes a `Reader` 
chunk by chunk and adds a bulk `read(byte[], int, int)`. `EncodingInputStream` 
now extends it; its API is unchanged.
   - **ftp, sftp, smb:** with `charset` set, 
`GenericFileHelper.toInputStream(exchange, charset)` reads the body as a 
`Reader` and encodes it while it is uploaded.
     - The `Reader` decodes the same way as the previous `String` conversion, 
so the written bytes are unchanged.
     - A `String` body, or a body with no `Reader` conversion, still takes the 
previous path.
   
   ## Deliberately not changed
   
   - **camel-http `disableStreamCache=true`.** The JIRA originally proposed 
that it also disable stream caching for the rest of the exchange. That would 
revert the intent of CAMEL-21162: the producer hands the raw response stream to 
the route's own stream caching, which spools to disk when enabled. 
`HttpStreamCacheSpoolToDiskTest` and `HttpStreamCacheNoSpoolToDiskTest` rely on 
this. To pass the response through without caching, use `.streamCache("false")` 
on the route.
   - **camel-vertx-http.** The hang with a remote file body is fixed separately 
by CAMEL-25154.
   - **google-storage** still stores `Content-Length` in the object's custom 
metadata, because existing tests assert it.
   - **The `CamelFileLength` header is not used as a length source.** It can be 
stale once the route changed the body, and a wrong length would truncate the 
upload.
   - **Unknown-length uploads still need one full copy** (to disk when spooling 
is enabled). Chunked multipart upload without a known length, streaming 
downloads, and platform-http pass-through are left for follow-ups.
   
   ## Tests
   
   New tests. The four regression tests fail on `main` and pass with this 
change:
   
   | Test | Result on `main` |
   |---|---|
   | `AWS2S3ProducerPayloadLengthTest#uploadPathBodyWithoutContentLength` | 
`Resetting to invalid mark` |
   | `BlobStreamAndLengthTest` | `Resetting to invalid mark` |
   | `FileStreamAndLengthTest` (datalake) | `Inputstream does not support mark 
rest operations` |
   | `GenericFileHelperTest#shouldEncodeStreamBodyWithCharsetWhileReading` | 
new helper; checks the body is encoded while read, not loaded as a whole |
   
   The other new tests pass:
   - 
`AWS2S3ProducerPayloadLengthTest#multiPartUploadStreamBodyWithUnknownLength`: 
the unknown-length fallback in the multipart path.
   - `PayloadHelperTest` (camel-core): length without reading (including a 
`BufferedInputStream` left unconsumed), wrapped-file length, and the fallback 
copy spooled to a file.
   - `SftpProducerWithCharsetIT#testProducerWithCharsetFromStream`: a stream 
body with `charset` goes through the new path.
   
   Existing tests run locally, all green:
   
   | Module | Tests |
   |---|---|
   | camel-util | 288 |
   | camel-core (charset, file, stream caching, XML converter, splitter suites) 
| 433 |
   | camel-file | 23 |
   | aws2-s3 | 78 |
   | azure-storage-blob | 58 |
   | azure-storage-datalake | 16 |
   | google-storage | 30 |
   | ibm-cos | 2 |
   | minio | 3 |
   | smb | 2 |
   | ftp/sftp producer, charset and stream-download ITs (embedded servers) | 84 
|
   | smb producer ITs (Docker) | 22 |
   | S3 upload/multipart ITs (LocalStack) | 9 |
   
   The MinIO ITs are `@Disabled` upstream, so MinIO relies on unit coverage.
   
   _Claude Code on behalf of Federico Mariani (Croway)_
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to