Croway opened a new pull request, #27132: URL: https://github.com/apache/camel/pull/27132
[CAMEL-25167](https://issues.apache.org/jira/browse/CAMEL-25167) ## Problem Object-storage producers each have their own way to find the length of a body that is not a `java.io.File`, and most of them either fail or load the whole body into memory: - **aws2-s3 and azure-storage-blob** probe the length with `mark(1024)` / `skip(available())` / `reset()`. On a `BufferedInputStream` past the mark buffer (about 8 KB) this throws `IOException: Resetting to invalid mark`. A `java.nio.file.Path` body is converted to exactly such a stream (`IOConverter.toInputStream(Path)`). For example, a Spring Boot platform-http multipart upload (the body is a `Path`) sent to `aws2-s3` fails for anything over a few KB. - **azure-storage-datalake** rejects streams without mark support. For the others it reads the whole stream into a `byte[]` to measure it. - **aws2-s3, minio, google-storage and ibm-cos** read a body of unknown length fully into a `ByteArrayOutputStream`. That bypasses stream caching, so spooling to disk never applies. The ftp, sftp and smb producers with `charset` set do `getMandatoryBody(String.class).getBytes(charset)`, which holds the whole file in memory twice. ## Changes - **camel-support:** new `PayloadHelper`. - `getBodyLength(Message)` / `getLength(Object)` return the length without reading the body: `WrappedFile` length, `File` / `Path` size, `StreamCache.length()`, `byte[]`, `ByteBuffer`, `ByteArrayInputStream`, `FileInputStream`. Otherwise they return `-1`. There is no mark/skip/reset probing. - `cacheStream(Exchange, InputStream)` copies a stream of unknown length through `OutputStreamBuilder`, so it is spooled to disk when stream-caching spooling is enabled. With spooling off, it behaves as before (in memory). - **aws2-s3, azure-storage-blob, azure-storage-datalake, minio, ibm-cos, google-storage:** use `PayloadHelper` instead of the per-component probes and in-memory copies. - A `Path` body now has a known length, so it is uploaded directly. - azure-storage-datalake accepts streams without mark support. - azure blob and datalake now convert a remote wrapped file (e.g. from SFTP) to a stream; before, they tried to convert the remote file handle. - `AWS2S3Utils.determineLengthInputStream`, `BlobUtils.getInputStreamLength` and `DataLakeUtils.getInputStreamLength` stay public and delegate to `PayloadHelper`. - **camel-util:** new `IOHelper.ReaderInputStream`, which encodes a `Reader` chunk by chunk and adds a bulk `read(byte[], int, int)`. `EncodingInputStream` now extends it; its API is unchanged. - **ftp, sftp, smb:** with `charset` set, `GenericFileHelper.toInputStream(exchange, charset)` reads the body as a `Reader` and encodes it while it is uploaded. - The `Reader` decodes the same way as the previous `String` conversion, so the written bytes are unchanged. - A `String` body, or a body with no `Reader` conversion, still takes the previous path. ## Deliberately not changed - **camel-http `disableStreamCache=true`.** The JIRA originally proposed that it also disable stream caching for the rest of the exchange. That would revert the intent of CAMEL-21162: the producer hands the raw response stream to the route's own stream caching, which spools to disk when enabled. `HttpStreamCacheSpoolToDiskTest` and `HttpStreamCacheNoSpoolToDiskTest` rely on this. To pass the response through without caching, use `.streamCache("false")` on the route. - **camel-vertx-http.** The hang with a remote file body is fixed separately by CAMEL-25154. - **google-storage** still stores `Content-Length` in the object's custom metadata, because existing tests assert it. - **The `CamelFileLength` header is not used as a length source.** It can be stale once the route changed the body, and a wrong length would truncate the upload. - **Unknown-length uploads still need one full copy** (to disk when spooling is enabled). Chunked multipart upload without a known length, streaming downloads, and platform-http pass-through are left for follow-ups. ## Tests New tests. The four regression tests fail on `main` and pass with this change: | Test | Result on `main` | |---|---| | `AWS2S3ProducerPayloadLengthTest#uploadPathBodyWithoutContentLength` | `Resetting to invalid mark` | | `BlobStreamAndLengthTest` | `Resetting to invalid mark` | | `FileStreamAndLengthTest` (datalake) | `Inputstream does not support mark rest operations` | | `GenericFileHelperTest#shouldEncodeStreamBodyWithCharsetWhileReading` | new helper; checks the body is encoded while read, not loaded as a whole | The other new tests pass: - `AWS2S3ProducerPayloadLengthTest#multiPartUploadStreamBodyWithUnknownLength`: the unknown-length fallback in the multipart path. - `PayloadHelperTest` (camel-core): length without reading (including a `BufferedInputStream` left unconsumed), wrapped-file length, and the fallback copy spooled to a file. - `SftpProducerWithCharsetIT#testProducerWithCharsetFromStream`: a stream body with `charset` goes through the new path. Existing tests run locally, all green: | Module | Tests | |---|---| | camel-util | 288 | | camel-core (charset, file, stream caching, XML converter, splitter suites) | 433 | | camel-file | 23 | | aws2-s3 | 78 | | azure-storage-blob | 58 | | azure-storage-datalake | 16 | | google-storage | 30 | | ibm-cos | 2 | | minio | 3 | | smb | 2 | | ftp/sftp producer, charset and stream-download ITs (embedded servers) | 84 | | smb producer ITs (Docker) | 22 | | S3 upload/multipart ITs (LocalStack) | 9 | The MinIO ITs are `@Disabled` upstream, so MinIO relies on unit coverage. _Claude Code on behalf of Federico Mariani (Croway)_ -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
