SEZ9 commented on issue #12295: URL: https://github.com/apache/seatunnel/issues/12295#issuecomment-5658990623
Thanks for the detailed report — reading the whole file via `Files.readAllBytes` and then materialising it again as a `String` is a real risk for long-running streaming jobs. A PR is welcome. A few requests for the PR: - Cover both the `LogBaseServlet#prepareLogResponse` path and the REST v1 copy in `RestHttpGetCommandProcessor`, ideally by deduplicating them into a shared helper, and keep the existing path-traversal guard unchanged. - Read only the tail rather than loading the whole file, and align the start to a line boundary as you propose so multi-byte UTF-8 characters are not split. - Consider making it visible to callers (e.g. a marker line or response header) that the content was truncated. - Document the new `seatunnel.engine.http.log-response-max-size-mb` option and add a unit test with a file larger than the limit. Could you also state the intended default value explicitly in the PR description? A large default that acts as a safety net without changing behaviour for typical files sounds like the right call. <!-- streview-comment:1038 --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
