[ 
https://issues.apache.org/jira/browse/CAMEL-25027?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Jiri Ondrusek updated CAMEL-25027:
----------------------------------
    Description: 
The langchain4j-ingest producer reads the body as text, splits it and embeds 
the segments. Media files cannot be ingested as media embeddings; they are 
decoded as garbled text and embedded anyway.
  
LangChain4j 1.20 added an experimental multimodal EmbeddingRequest API: an 
input can be image, audio, video or PDF content, and a model declares what it 
accepts in supportedContentTypes().
* Add a modality endpoint option (text, the default, or media). With 
modality=media: the body is read as bytes and embedded whole, as one vector
* the vector is stored under a placeholder TextSegment whose text is the 
document id, with the same identity metadata as text segments
* the MIME type comes from a new contentType option or from the document id's 
extension; headers are not consulted
* the endpoint refuses to start with a model that declares no media type, with 
a documentSplitter, or with contentType under modality=text; a document whose 
medium the model lacks fails before its dedup claim and body read
* maxDocumentSize and minDocumentSize count bytes; splitter and batch options 
do not apply
* filtering, deduplication and results are unchanged

Image is the only medium a LangChain4j provider embeds today (Jina, Voyage AI, 
Cohere, Gemini, Bedrock Titan). Audio, video and PDF need a custom 
EmbeddingModel.

  was:
The langchain4j-ingest producer only handles text: the body is read as a 
String, split into segments and embedded with embedAll. Audio files cannot be 
ingested as audio
  embeddings, the way a Wav2Vec2 or CLAP model would embed them for acoustic 
similarity search. Without a parser they are decoded as garbage text and 
embedded anyway.

  LangChain4j 1.19 has the API for this: EmbeddingRequest accepts AudioContent 
and a model declares AUDIO in supportedContentTypes().
 
  Proposal: a modality endpoint option, text (default, unchanged behaviour) or 
audio.

  With modality=audio:
  - the body is read as bytes and embedded whole, as one vector, through 
embed(EmbeddingRequest) with an AudioContent
  - the vector is stored with a placeholder TextSegment whose text is the 
document id, carrying the same camel_ingest_pipeline and 
camel_ingest_document_id metadata as text
    segments
  - the endpoint fails to start when the model does not declare AUDIO
  - the MIME type comes from a new contentType option or, when unset, from the 
document id's file extension (wav, mp3, flac, ogg, m4a, aac); an unknown type 
fails the
    exchange
  - maxDocumentSize and minDocumentSize count bytes; maxSegmentSize, 
maxOverlapSize and embeddingBatchSize do not apply; documentSplitter is rejected
  - id filters, dedup claim and release, documentFilter and IngestResult 
outcomes work unchanged

  Out of scope: chunking long audio into time windows, storing the MIME type as 
metadata, sniffing the type from the bytes.

  Note: no LangChain4j provider embeds audio out of the box yet, so the model 
is typically a custom EmbeddingModel wrapping an audio embedding service. Tests 
use a
  deterministic audio-capable fake.

  Follow-ups planned in camel-quarkus (extension option on camel-main) and 
camel-kamelets (langchain4j-ingest-sink property).



> camel-langchain4j-ingest - Add media modality to embed images, audio, video 
> and PDF whole   
> --------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25027
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25027
>             Project: Camel
>          Issue Type: New Feature
>          Components: camel-langchain4j
>    Affects Versions: 4.23.0
>            Reporter: Jiri Ondrusek
>            Assignee: Jiri Ondrusek
>            Priority: Major
>
> The langchain4j-ingest producer reads the body as text, splits it and embeds 
> the segments. Media files cannot be ingested as media embeddings; they are 
> decoded as garbled text and embedded anyway.
>   
> LangChain4j 1.20 added an experimental multimodal EmbeddingRequest API: an 
> input can be image, audio, video or PDF content, and a model declares what it 
> accepts in supportedContentTypes().
> * Add a modality endpoint option (text, the default, or media). With 
> modality=media: the body is read as bytes and embedded whole, as one vector
> * the vector is stored under a placeholder TextSegment whose text is the 
> document id, with the same identity metadata as text segments
> * the MIME type comes from a new contentType option or from the document id's 
> extension; headers are not consulted
> * the endpoint refuses to start with a model that declares no media type, 
> with a documentSplitter, or with contentType under modality=text; a document 
> whose medium the model lacks fails before its dedup claim and body read
> * maxDocumentSize and minDocumentSize count bytes; splitter and batch options 
> do not apply
> * filtering, deduplication and results are unchanged
> Image is the only medium a LangChain4j provider embeds today (Jina, Voyage 
> AI, Cohere, Gemini, Bedrock Titan). Audio, video and PDF need a custom 
> EmbeddingModel.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to