SEZ9 commented on issue #12627: URL: https://github.com/apache/seatunnel/issues/12627#issuecomment-6030554977
Thanks for the detailed report — the gap you isolated is real: the repair path head-truncates the failure text at 3,000 characters, so for a long Java trace the trailing `Caused by` chain and the `ErrorCode` are exactly the parts that get dropped before the model ever sees them, while the local tail capture still shows them. On top of that, the `parse_error()` helper added earlier is not actually called on the production path, so the structured code/description was never reaching the repair prompt. The fix approach being reviewed looks right to me: parse the complete failure first so the error code and description survive regardless of length, then shorten the raw excerpt by keeping the tail of an oversized trace (where the innermost `Caused by` lives) instead of the head, and redact both the parsed fields and the excerpt before they go into the prompt. Covering the code/root cause beyond the 3,000-character boundary, tail-vs-head selection, and the exact-limit and short-trace controls in regression tests is the right set of cases. Remaining asks: - Please keep this issue open until the fix PR is merged into `dev`; we don't want a second implementation path for the same change. - On the PR side, I'd like to see confirmation that `parse_error()` is wired into the actual CLI repair flow (not only exercised by tests), so the parsed `ErrorCode`/description is what the model receives. - Once it lands, it would be great if you could re-run your original failing job through the AI CLI and confirm the root cause and error code now appear in the diagnosis. <!-- streview-comment:1579 --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
