[
https://issues.apache.org/jira/browse/CAMEL-25027?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Jiri Ondrusek updated CAMEL-25027:
----------------------------------
Description:
The langchain4j-ingest producer reads the body as text, splits it and embeds
the segments. Media files cannot be ingested as media embeddings; they are
decoded as garbled text and embedded anyway.
LangChain4j 1.20 added an experimental multimodal EmbeddingRequest API: an
input can be image, audio, video or PDF content, and a model declares what it
accepts in supportedContentTypes().
* Add a modality endpoint option (text, the default, or media). With
modality=media: the body is read as bytes and embedded whole, as one vector
* the vector is stored under a placeholder TextSegment whose text is the
document id, with the same identity metadata as text segments
* the MIME type comes from a new contentType option or from the document id's
extension; headers are not consulted
* the endpoint refuses to start with a model that declares no media type, with
a documentSplitter, or with contentType under modality=text; a document whose
medium the model lacks fails before its dedup claim and body read
* maxDocumentSize and minDocumentSize count bytes; splitter and batch options
do not apply
* filtering, deduplication and results are unchanged
Image is the only medium a LangChain4j provider embeds today (Jina, Voyage AI,
Cohere, Gemini, Bedrock Titan). Audio, video and PDF need a custom
EmbeddingModel.
was:
The langchain4j-ingest producer only handles text: the body is read as a
String, split into segments and embedded with embedAll. Audio files cannot be
ingested as audio
embeddings, the way a Wav2Vec2 or CLAP model would embed them for acoustic
similarity search. Without a parser they are decoded as garbage text and
embedded anyway.
LangChain4j 1.19 has the API for this: EmbeddingRequest accepts AudioContent
and a model declares AUDIO in supportedContentTypes().
Proposal: a modality endpoint option, text (default, unchanged behaviour) or
audio.
With modality=audio:
- the body is read as bytes and embedded whole, as one vector, through
embed(EmbeddingRequest) with an AudioContent
- the vector is stored with a placeholder TextSegment whose text is the
document id, carrying the same camel_ingest_pipeline and
camel_ingest_document_id metadata as text
segments
- the endpoint fails to start when the model does not declare AUDIO
- the MIME type comes from a new contentType option or, when unset, from the
document id's file extension (wav, mp3, flac, ogg, m4a, aac); an unknown type
fails the
exchange
- maxDocumentSize and minDocumentSize count bytes; maxSegmentSize,
maxOverlapSize and embeddingBatchSize do not apply; documentSplitter is rejected
- id filters, dedup claim and release, documentFilter and IngestResult
outcomes work unchanged
Out of scope: chunking long audio into time windows, storing the MIME type as
metadata, sniffing the type from the bytes.
Note: no LangChain4j provider embeds audio out of the box yet, so the model
is typically a custom EmbeddingModel wrapping an audio embedding service. Tests
use a
deterministic audio-capable fake.
Follow-ups planned in camel-quarkus (extension option on camel-main) and
camel-kamelets (langchain4j-ingest-sink property).
> camel-langchain4j-ingest - Add media modality to embed images, audio, video
> and PDF whole
> --------------------------------------------------------------------------------------------
>
> Key: CAMEL-25027
> URL: https://issues.apache.org/jira/browse/CAMEL-25027
> Project: Camel
> Issue Type: New Feature
> Components: camel-langchain4j
> Affects Versions: 4.23.0
> Reporter: Jiri Ondrusek
> Assignee: Jiri Ondrusek
> Priority: Major
>
> The langchain4j-ingest producer reads the body as text, splits it and embeds
> the segments. Media files cannot be ingested as media embeddings; they are
> decoded as garbled text and embedded anyway.
>
> LangChain4j 1.20 added an experimental multimodal EmbeddingRequest API: an
> input can be image, audio, video or PDF content, and a model declares what it
> accepts in supportedContentTypes().
> * Add a modality endpoint option (text, the default, or media). With
> modality=media: the body is read as bytes and embedded whole, as one vector
> * the vector is stored under a placeholder TextSegment whose text is the
> document id, with the same identity metadata as text segments
> * the MIME type comes from a new contentType option or from the document id's
> extension; headers are not consulted
> * the endpoint refuses to start with a model that declares no media type,
> with a documentSplitter, or with contentType under modality=text; a document
> whose medium the model lacks fails before its dedup claim and body read
> * maxDocumentSize and minDocumentSize count bytes; splitter and batch options
> do not apply
> * filtering, deduplication and results are unchanged
> Image is the only medium a LangChain4j provider embeds today (Jina, Voyage
> AI, Cohere, Gemini, Bedrock Titan). Audio, video and PDF need a custom
> EmbeddingModel.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)