+1 for the proposal. Looking forward to Paimon 2.0! Best, wangwj
On Tue, Jul 7, 2026 at 3:24 PM Yunfeng Zhou <[email protected]> wrote: > > Hi Jingsong, > > +1 for the proposal. Looking forward to Paimon 2.0! > > Best, > Yunfeng > > > 2026年7月7日 11:11,Jingsong Li <[email protected]> 写道: > > > > Hi everyone, > > > > I would like to start a discussion about preparing an Apache Paimon > > 2.0.0 release. > > > > Over the past release cycles, Paimon has evolved significantly. In > > addition to continuing improvements to the core lakehouse table > > engine, we have been building a broader storage foundation for > > streaming, analytics, and AI/multimodal workloads. > > > > Some major areas that have landed or are being stabilized include: > > > > 1. Data Evolution and row tracking > > > > Paimon now has a much stronger foundation for append tables with row > > tracking and data evolution. This enables efficient partial column > > updates, schema evolution, dedicated storage for special column types, > > and global-index-based retrieval without rewriting entire data files. > > > > 2. Multimodal table capabilities > > > > We have introduced multimodal table support, covering blob storage, > > vector storage, full-text content, and global indexes in one table > > abstraction. This allows Paimon tables to store and query structured > > data together with images, videos, audio, embeddings, and text > > content. > > > > 3. Global Index framework > > > > The global index framework has been extended to support multiple index > > types, including BTree, Bitmap, vector indexes, full-text indexes, and > > hybrid search. This provides a unified mechanism for scalar filtering, > > vector similarity search, text retrieval, and combined retrieval > > workflows. > > > > 4. Vector and full-text search > > > > We have added native vector index support, including paimon-vindex > > IVF-based indexes and Lumina DiskANN-based indexes, as well as native > > full-text search integration with BM25 scoring and tokenizer > > configuration. These features are important for RAG, recommendation, > > image retrieval, and other AI-native scenarios. > > > > 5. New storage formats and supporting components > > > > We have also made several supporting components available around > > Paimon, including Mosaic, paimon-vindex, and paimon-full-text-index. > > > > Mosaic provides a columnar-bucket hybrid file format optimized for > > wide tables and efficient column projection. paimon-vindex provides > > native vector indexing capabilities for Paimon. paimon-full-text-index > > provides the native full-text engine used by Paimon full-text global > > indexes. > > > > Together, these components make Paimon more capable as a storage layer > > for both traditional lakehouse workloads and emerging AI/multimodal > > workloads. > > > > 6. Ecosystem improvements > > > > We have continued to improve Flink, Spark, Python, Ray, and Daft > > integration, and have been adding more end-to-end coverage for > > Java/Python interoperability, global index creation and reading, > > vector data, full-text search, and data evolution workflows. > > > > Given the scope of these changes, I think it is a good time to discuss > > whether we should cut a 2.0.0 release. A 2.0 release would help > > communicate that Paimon has entered a new stage, with data evolution, > > multimodal storage, and global indexes becoming first-class parts of > > the project. > > > > Please share your thoughts, concerns, or any issues that you believe > > should block or be included in the 2.0.0 release. > > > > Best, > > Jingsong >
