+1 for the proposal. Looking forward to Paimon 2.0!

Best,
wangwj

On Tue, Jul 7, 2026 at 3:24 PM Yunfeng Zhou <[email protected]> wrote:
>
> Hi Jingsong,
>
> +1 for the proposal. Looking forward to Paimon 2.0!
>
> Best,
> Yunfeng
>
> > 2026年7月7日 11:11,Jingsong Li <[email protected]> 写道:
> >
> > Hi everyone,
> >
> > I would like to start a discussion about preparing an Apache Paimon
> > 2.0.0 release.
> >
> > Over the past release cycles, Paimon has evolved significantly. In
> > addition to continuing improvements to the core lakehouse table
> > engine, we have been building a broader storage foundation for
> > streaming, analytics, and AI/multimodal workloads.
> >
> > Some major areas that have landed or are being stabilized include:
> >
> > 1. Data Evolution and row tracking
> >
> > Paimon now has a much stronger foundation for append tables with row
> > tracking and data evolution. This enables efficient partial column
> > updates, schema evolution, dedicated storage for special column types,
> > and global-index-based retrieval without rewriting entire data files.
> >
> > 2. Multimodal table capabilities
> >
> > We have introduced multimodal table support, covering blob storage,
> > vector storage, full-text content, and global indexes in one table
> > abstraction. This allows Paimon tables to store and query structured
> > data together with images, videos, audio, embeddings, and text
> > content.
> >
> > 3. Global Index framework
> >
> > The global index framework has been extended to support multiple index
> > types, including BTree, Bitmap, vector indexes, full-text indexes, and
> > hybrid search. This provides a unified mechanism for scalar filtering,
> > vector similarity search, text retrieval, and combined retrieval
> > workflows.
> >
> > 4. Vector and full-text search
> >
> > We have added native vector index support, including paimon-vindex
> > IVF-based indexes and Lumina DiskANN-based indexes, as well as native
> > full-text search integration with BM25 scoring and tokenizer
> > configuration. These features are important for RAG, recommendation,
> > image retrieval, and other AI-native scenarios.
> >
> > 5. New storage formats and supporting components
> >
> > We have also made several supporting components available around
> > Paimon, including Mosaic, paimon-vindex, and paimon-full-text-index.
> >
> > Mosaic provides a columnar-bucket hybrid file format optimized for
> > wide tables and efficient column projection. paimon-vindex provides
> > native vector indexing capabilities for Paimon. paimon-full-text-index
> > provides the native full-text engine used by Paimon full-text global
> > indexes.
> >
> > Together, these components make Paimon more capable as a storage layer
> > for both traditional lakehouse workloads and emerging AI/multimodal
> > workloads.
> >
> > 6. Ecosystem improvements
> >
> > We have continued to improve Flink, Spark, Python, Ray, and Daft
> > integration, and have been adding more end-to-end coverage for
> > Java/Python interoperability, global index creation and reading,
> > vector data, full-text search, and data evolution workflows.
> >
> > Given the scope of these changes, I think it is a good time to discuss
> > whether we should cut a 2.0.0 release. A 2.0 release would help
> > communicate that Paimon has entered a new stage, with data evolution,
> > multimodal storage, and global indexes becoming first-class parts of
> > the project.
> >
> > Please share your thoughts, concerns, or any issues that you believe
> > should block or be included in the 2.0.0 release.
> >
> > Best,
> > Jingsong
>

Reply via email to