SEZ9 commented on issue #12295:
URL: https://github.com/apache/seatunnel/issues/12295#issuecomment-5658990623

   Thanks for the detailed report — reading the whole file via 
`Files.readAllBytes` and then materialising it again as a `String` is a real 
risk for long-running streaming jobs. A PR is welcome.
   
   A few requests for the PR:
   
   - Cover both the `LogBaseServlet#prepareLogResponse` path and the REST v1 
copy in `RestHttpGetCommandProcessor`, ideally by deduplicating them into a 
shared helper, and keep the existing path-traversal guard unchanged.
   - Read only the tail rather than loading the whole file, and align the start 
to a line boundary as you propose so multi-byte UTF-8 characters are not split.
   - Consider making it visible to callers (e.g. a marker line or response 
header) that the content was truncated.
   - Document the new `seatunnel.engine.http.log-response-max-size-mb` option 
and add a unit test with a file larger than the limit.
   
   Could you also state the intended default value explicitly in the PR 
description? A large default that acts as a safety net without changing 
behaviour for typical files sounds like the right call.
   
   <!-- streview-comment:1038 -->


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to