[ 
https://issues.apache.org/jira/browse/NIFI-16195?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Daniel Stieglitz reassigned NIFI-16195:
---------------------------------------

    Assignee: Daniel Stieglitz

> ExtractEmailAttachments: Write part headers into attributes
> -----------------------------------------------------------
>
>                 Key: NIFI-16195
>                 URL: https://issues.apache.org/jira/browse/NIFI-16195
>             Project: Apache NiFi
>          Issue Type: Improvement
>          Components: Extensions
>    Affects Versions: 2.11.0
>         Environment: N/A
>            Reporter: René Zeidler
>            Assignee: Daniel Stieglitz
>            Priority: Major
>              Labels: attachments, email
>
> The {{ExtractEmailAttachments}} processor creates a FlowFile for each "part" 
> of an email. For an HTML email, this can be many different parts:
>  * The HTML body
>  * An alternative plain text representation of the body
>  * Inline images
>  * Actual attachments
> Currently there is no easy way to filter out the parts you actually want 
> (e.g. just the attachments but not the inline images) except guessing from 
> the content of the attachment. Each part has a Content-Type header, and 
> optional Content-ID, Content-Description, Content-Disposition, and 
> Content-Transfer-Encoding headers.
> This is the structure of an example email:
>  * Content-Type: multipart/mixed; boundary="...1..."
>  ** Content-Type: multipart/related; type="multipart/alternative"; 
> boundary="...2..."
>  *** Content-Type: multipart/alternative; boundary="...3..."
>  **** Content-Type: text/plain; charset="iso-8859-1"
> Content-Transfer-Encoding: quoted-printable
> _<1: plaintext body>_
>  **** Content-Type: text/html; charset="iso-8859-1"
> Content-Transfer-Encoding: quoted-printable
> _<2: html body>_
>  *** Content-ID: <0011223344556677@8899AABBCCDDEEFF>
> Content-Type: image/png; name="inline_image.png"
> Content-Disposition: inline; filename="inline_image.png"
> Content-Transfer-Encoding: base64
> _<3: inline image>_
>  ** Content-Type: application/pdf; name="my-attachment.pdf"
> Content-Description: my-attachment.pdf
> Content-Disposition: attachment;
>     filename="my-attachment.pdf"; size=71521;
>     creation-date="Thu, 13 Aug 2026 11:02:50 GMT";
>     modification-date="Thu, 13 Aug 2026 11:01:24 GMT"
> Content-Transfer-Encoding: base64
> _<4: pdf attachment>_
> The processor correctly returns all four leaf parts of the message, but 
> without a way to distinguish them. The simplest solution would be to write 
> all {{Content-*}} headers into attributes.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to