Il giorno mer 12 ago 2026 alle ore 15:34 tison <[email protected]> ha
scritto:

>
> https://cwiki.apache.org/confluence/spaces/INCUBATOR/pages/446071456/OpenCrawling+Proposal
>
> This page is ready. May you check if you have permission to edit this page?
> Note that you may log in with your Apache ID.
>

Unfortunately I can't edit the page with my Apache account.


>
>
> Best,
> tison.
>
>
> Piergiorgio Lucidi <[email protected]> 于2026年8月12日周三 21:31写道:
>
> > Hi tison,
> >
> > Thank you so much for publishing the proposal and for trying to solve
> this
> > but I'm continuing to have the same issue.
> > I tried to create a blank page giving Proposals as the main page but I
> see
> > the same permission error.
> >
> > Cheers,
> > PG
> >
> >
> >
> > Il giorno mer 12 ago 2026 alle ore 03:00 tison <[email protected]> ha
> > scritto:
> >
> > > Hi Piergiorgio,
> > >
> > > I've invited you to the page. But I'm unfamiliar with Confluence, so it
> > may
> > > not be what you're looking for.
> > >
> > > Alternatively, you may find
> > > https://selfserve.apache.org/confluence-account.html helps.
> > >
> > > Anyway, I created the proposal page [1] with the wiki page content you
> > > shared.
> > >
> > > [1]
> > >
> > >
> >
> https://cwiki.apache.org/confluence/spaces/INCUBATOR/pages/446071456/OpenCrawling+Proposal
> > >
> > > Best,
> > > tison.
> > >
> > >
> > > Piergiorgio Lucidi <[email protected]> 于2026年8月12日周三 03:36写道:
> > >
> > > > Hi PJ,
> > > >
> > > > Thank you for the feedback and for highlighting Apache StormCrawler.
> > > >
> > > > I completely agree that exploring collaboration between the two
> > projects
> > > is
> > > > a great idea.
> > > >
> > > > While StormCrawler is an incredibly powerful tool for large scale web
> > > > crawling and processing, OpenCrawling was built specifically to
> tackle
> > > the
> > > > enterprise content and process automation ecosystem.
> > > >
> > > > Our core focus is bridging platforms like Alfresco, Flowable, and
> > Camunda
> > > > directly into modern LLM and RAG architectures, heavily utilizing
> > Spring
> > > > Boot and Spring AI.
> > > >
> > > > Despite the distinct use cases, web data versus enterprise
> > repositories,
> > > > there is absolutely a shared interest in robust data ingestion,
> > document
> > > > parsing and vector database integration strategies.
> > > >
> > > > OpenCrawling can become the perfect home for any crawling strategy.
> > > >
> > > > I would be thrilled to connect with the StormCrawler community to see
> > how
> > > > our projects might complement each other and share best practices
> > moving
> > > > forward.
> > > >
> > > > We could propose to implement a brand new OpenCrawling Storm Bolt.
> > > >
> > > > In Apache Storm topology, data flows from *Spouts* (URL queues) to
> > > *Bolts*
> > > > (fetchers, parsers and indexers). The most seamless integration is to
> > > build
> > > > an opencrawling-storm-bolt.
> > > >
> > > > StormCrawler handles the heavy lifting of recursive web crawling,
> > > > politeness and HTML parsing.
> > > >
> > > > Instead of using StormCrawler's native OpenSearch/Elasticsearch
> indexer
> > > > bolt, the topology passes the parsed document to the new custom
> > > > OpenCrawling Storm Bolt.
> > > >
> > > > The Bolt acts as an OpenCrawling Repository Connector. It takes the
> raw
> > > > HTML, wraps it in the Open Ingestion Standard (OIS) format, attaches
> > any
> > > > relevant baseline metadata and pushes it through OpenCrawling’s
> secure
> > > > pipeline (Java 25 / Spring Boot 4) for chunking, embedding, and
> Vector
> > DB
> > > > ingestion.
> > > >
> > > > Cheers,
> > > >
> > > > PG
> > > >
> > > > Il Mar 11 Ago 2026, 18:44 PJ Fanning <[email protected]> ha
> > scritto:
> > > >
> > > > > This doesn't block OpenCrawling joining as an ASF Incubator podling
> > > > > but I just want to highlight that there is already Apache
> > > > > StormCrawler.
> > > > >
> > > > > https://stormcrawler.apache.org/
> > > > >
> > > > > It would be great if these projects could collaborate in areas of
> > > > > shared interest.
> > > > >
> > > > > On Tue, 11 Aug 2026 at 15:28, Piergiorgio Lucidi <
> > > [email protected]
> > > > >
> > > > > wrote:
> > > > > >
> > > > > > Hi everyone,
> > > > > >
> > > > > > I would like to propose OpenCrawling as a new project for
> > incubation
> > > > > within
> > > > > > the Apache Software Foundation. Currently the project is hosted
> on
> > > > GitHub
> > > > > > [1].
> > > > > >
> > > > > > The official proposal is currently available in the OpenCrawling
> > Wiki
> > > > in
> > > > > > markdown format [2].
> > > > > >
> > > > > > I tried to share the proposal in our Confluence but it seems
> that I
> > > > don't
> > > > > > have permission to create the new page under the Proposals page.
> > > Anyway
> > > > > if
> > > > > > someone can guide me on resolving this issue it would be great!
> > > > > >
> > > > > > Once space permissions are granted on cwiki.apache.org, I will
> > also
> > > > > mirror
> > > > > > the proposal on the Incubator CWIKI proposals page.
> > > > > > Below you also find the same proposal ready to be copy-pasted
> into
> > > our
> > > > > > Confluence.
> > > > > >
> > > > > > We welcome feedback, questions and discussion from the Incubator
> > > > > community!
> > > > > >
> > > > > > Best regards,
> > > > > > Piergiorgio
> > > > > > On behalf of the OpenCrawling Core Team
> > > > > >
> > > > > > [1] - https://github.com/opencrawling/opencrawling
> > > > > > [2] -
> > > > > >
> > > > >
> > > >
> > >
> >
> https://github.com/opencrawling/opencrawling/wiki/Apache-Incubator-Proposal
> > > > > >
> > > > > > --------------------------------------------------------------
> > > > > >
> > > > > > h1. Apache Incubator Proposal: OpenCrawling
> > > > > >
> > > > > > h2. Abstract
> > > > > >
> > > > > > *OpenCrawling* is an open-source, enterprise-grade,
> > high-performance
> > > > data
> > > > > > crawling, content ingestion, and security-aware vector search
> > > platform.
> > > > > > Built on modern Java 25 (leveraging Virtual Threads and
> Structured
> > > > > > Concurrency), Spring Boot 4, and Spring AI, OpenCrawling serves
> as
> > > the
> > > > > > reference implementation of the *Open Ingestion Standard (OIS)*
> and
> > > > > > provides a secure *Model Context Protocol (MCP)* server
> interface.
> > It
> > > > > > orchestrates scalable data flows from heterogeneous enterprise
> > > > > repositories
> > > > > > (e.g., SharePoint, S3, CMIS/Alfresco, BPMN engines like Camunda
> and
> > > > > > Flowable, Apache Iceberg, Apache Ozone) to downstream vector
> > > databases
> > > > > > (e.g., Milvus, Qdrant, OpenSearch, Vespa, pgvector) with
> > source-level
> > > > > > Access Control List (ACL) security enforcement.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Proposal
> > > > > >
> > > > > > The OpenCrawling community proposes to incubate *OpenCrawling*
> as a
> > > new
> > > > > > project within the Apache Software Foundation (ASF). OpenCrawling
> > > > > provides
> > > > > > a decoupled, vendor-neutral enterprise data integration framework
> > > that
> > > > > > bridges the gap between traditional enterprise content management
> > > (ECM)
> > > > > > repositories and modern Large Language Model (LLM) /
> > > > Retrieval-Augmented
> > > > > > Generation (RAG) architectures.
> > > > > >
> > > > > > The project encompasses:
> > > > > > # *Core Ingestion Runtime ({{oc-core}}, {{oc-runtime}})*: A
> > > > distributed,
> > > > > > asynchronous engine built with virtual threads, claim-check
> > metadata
> > > > > > patterns, and Apache Kafka event streams.
> > > > > > # *Repository & Vector Connectors*: Standardized connectors for
> > > > scanning
> > > > > > source systems and indexing vector embeddings into major vector
> > > stores.
> > > > > > # *Open Ingestion Standard (OIS)*: The formal JSON/YAML schema
> > > > > > specifications defining unified document payloads, ACL security
> > SIDs,
> > > > and
> > > > > > crawler job configurations.
> > > > > > # *Secure Model Context Protocol (MCP) Server*: A Zero-Trust
> > context
> > > > > > retrieval server enforcing document-level permissions (Active
> > > Directory
> > > > > > SIDs, LDAP groups, user principals) at query time.
> > > > > > # *Observability & Developer Tooling*: AI-Powered Observability
> > > (AIOps)
> > > > > > over OpenTelemetry traces, Auto-Narrativization Copilot, Java
> > Client
> > > > SDK,
> > > > > > and Maven Archetypes for custom connector development.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Background
> > > > > >
> > > > > > In enterprise AI and RAG architectures, LLM agents require
> seamless
> > > > > access
> > > > > > to unstructured content stored across legacy and cloud
> > repositories.
> > > > > > However, traditional ingestion pipelines often strip out or
> ignore
> > > > > > source-level security metadata (ACLs), leading to context leakage
> > > where
> > > > > an
> > > > > > AI model synthesizes responses using confidential documents that
> > the
> > > > > > requesting user does not have permissions to view.
> > > > > >
> > > > > > Furthermore, legacy crawling tools (such as Apache ManifoldCF or
> > > Apache
> > > > > > Nutch) were architected over a decade ago prior to the emergence
> of
> > > > > vector
> > > > > > databases, LLMs, Model Context Protocol (MCP), and modern Java
> > > features
> > > > > > like Virtual Threads (JEP 444) and Structured Concurrency.
> > > > > >
> > > > > > OpenCrawling was created to address this modern ingestion crisis
> by
> > > > > > providing a native Java 25/Spring AI implementation engineered
> > > > > specifically
> > > > > > for LLM search scenarios, zero-trust context retrieval, and
> > > > > high-throughput
> > > > > > asynchronous processing.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Rationale
> > > > > >
> > > > > > The Apache Software Foundation is the natural home for
> > OpenCrawling.
> > > > ASF
> > > > > > has long been the center of innovation for enterprise search and
> > big
> > > > data
> > > > > > infrastructure, hosting cornerstone projects such as Apache
> Lucene,
> > > > > Apache
> > > > > > Solr, Apache Tika, Apache Kafka, Apache Iceberg, Apache Ozone,
> > Apache
> > > > > > ManifoldCF, and Apache Nutch.
> > > > > >
> > > > > > Bringing OpenCrawling to the ASF offers multiple mutual benefits:
> > > > > > * *Ecosystem Integration*: OpenCrawling directly integrates with
> > and
> > > > > builds
> > > > > > upon existing Apache projects, including *Apache Tika* (text
> > > > extraction),
> > > > > > *Apache Kafka* (event-driven pipeline), *Apache Iceberg*
> (lakehouse
> > > > > > connector), *Apache Ozone* (claim-check object storage), and
> > *Apache
> > > > > Maven*
> > > > > > (connector archetype distribution).
> > > > > > * *Vendor-Neutral Governance*: Neutral governance under the
> Apache
> > > Way
> > > > is
> > > > > > vital to establishing OpenCrawling and OIS as industry-wide,
> > > > > > vendor-agnostic ingestion standards.
> > > > > > * *Community Sustainability*: Operating as an Apache project will
> > > > > attract a
> > > > > > broader community of enterprise adopters, cloud providers, AI
> > > framework
> > > > > > developers, and search engine vendors.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Initial Goals
> > > > > >
> > > > > > During incubation, the OpenCrawling project will focus on the
> > > following
> > > > > > milestones:
> > > > > >
> > > > > > # *ASF Migration & Infrastructure*:
> > > > > > ** Transfer codebases ({{opencrawling}},
> > {{open-ingestion-standard}},
> > > > {{
> > > > > > opencrawling.github.io}}) to Apache infrastructure ({{
> > > > > > github.com/apache/incubator-opencrawling}}
> > > > > <http://github.com/apache/incubator-opencrawling%7D%7D>).
> > > > > > ** Rebrand build artifacts to {{org.apache.opencrawling}}.
> > > > > > ** Setup ASF-compliant CI/CD pipelines using GitHub Actions.
> > > > > > # *Community & Governance*:
> > > > > > ** Adopt the Apache Way for all decisions, roadmap discussions,
> and
> > > > > release
> > > > > > voting.
> > > > > > ** Expand the contributor base across independent developers,
> > > > enterprise
> > > > > > search users, and corporate contributors.
> > > > > > # *Ecosystem & Connector Expansion*:
> > > > > > ** Release additional output connectors (Elasticsearch, Apache
> > Solr,
> > > > > > RESTHeart).
> > > > > > ** Add native integration for fine-grained authorization
> frameworks
> > > > > (e.g.,
> > > > > > OpenFGA).
> > > > > > ** Enhance gRPC support for high-efficiency inter-microservice
> > > > > > communication.
> > > > > > ** Standardize OIS specification drafts under ASF governance.
> > > > > > # *Compliance & Licensing*:
> > > > > > ** Complete IP clearance and execute software grant agreements.
> > > > > > ** Ensure all third-party dependencies strictly conform to Apache
> > > > License
> > > > > > Category A policies.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Current Status
> > > > > >
> > > > > > h3. Meritocracy
> > > > > > The OpenCrawling project was established with meritocratic
> > principles
> > > > > from
> > > > > > day one. Design decisions, architecture changes, issue tracking,
> > and
> > > > > > roadmap discussions take place openly on GitHub through RFCs,
> Pull
> > > > > > Requests, and public wiki pages.
> > > > > >
> > > > > > h3. Community
> > > > > > The OpenCrawling community includes developers and architects
> from
> > > > > > enterprise search, ECM, and AI background. Community channels
> > include
> > > > > > GitHub Discussions, Slack, and social media announcements. The
> > > project
> > > > > > actively encourages external contributions via Maven archetypes
> and
> > > > > modular
> > > > > > connector development.
> > > > > >
> > > > > > h3. Core Developers
> > > > > > The initial core developers are experienced software architects
> and
> > > > > > open-source veterans with extensive experience in enterprise
> > search,
> > > > > > content management, and ASF governance:
> > > > > >
> > > > > > * *Piergiorgio Lucidi* ({{[email protected]}}) – Founder,
> > Lead
> > > > > > Architect. ASF Member and PMC Member/Committer on multiple Apache
> > > > > projects
> > > > > > (including Apache ManifoldCF and Apache Chemistry).
> > > > > > * *Michael Cizmar* ({{[email protected]}}) – Lead
> > Architect
> > > &
> > > > > > Developer. Specialist in enterprise search and cloud
> > infrastructure.
> > > > > > * *Luis Cabaceira* ({{[email protected]}}) – Lead
> > Architect &
> > > > > > Developer. Specialist in document processing and AI integration.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Known Risks
> > > > > >
> > > > > > h3. Orphaned Products
> > > > > > The risk of OpenCrawling becoming orphaned is low. The project
> > solves
> > > > an
> > > > > > active, urgent security and performance problem in enterprise AI
> > > > adoption
> > > > > > (RAG ACL context leakage). The core maintainers are committed to
> > its
> > > > > > long-term evolution and actively use it in production
> environments.
> > > > > >
> > > > > > h3. Inexperience with Open Source
> > > > > > The project leadership has deep experience with open-source
> > > > communities.
> > > > > > Piergiorgio Lucidi is an active ASF Member and PMC member with
> > over a
> > > > > > decade of experience guiding projects through the Apache Way.
> > > > > >
> > > > > > h3. Homogenous Developers
> > > > > > The initial committers come from diverse geographical locations
> > > (Italy,
> > > > > > United States, Portugal) and distinct
> organizations/consultancies.
> > > > > > Incubating at Apache will further diversify the developer base by
> > > > > > encouraging contributions from enterprise organizations and
> search
> > > > > vendors.
> > > > > >
> > > > > > h3. Reliance on Third-Party Products
> > > > > > OpenCrawling is designed to be vendor-neutral. Core dependencies
> > are
> > > > > > open-source libraries under permissive licenses (Apache 2.0, MIT,
> > > BSD):
> > > > > > * Spring Boot & Spring AI (Apache 2.0)
> > > > > > * Apache Tika (Apache 2.0)
> > > > > > * Apache Kafka (Apache 2.0)
> > > > > > * PostgreSQL / pgvector (PostgreSQL License / MIT)
> > > > > > * Docker & OpenTelemetry (Apache 2.0)
> > > > > >
> > > > > > There are no GPL/AGPL dependencies in the runtime core.
> > > > > >
> > > > > > h3. Relationship with Sponsored Products / Brand
> > > > > > OpenCrawling is an independent project. The name "OpenCrawling"
> has
> > > > been
> > > > > > used for the open-source codebase. The trademark will be
> > transferred
> > > to
> > > > > the
> > > > > > Apache Software Foundation upon incubation acceptance.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Documentation & Existing Artifacts
> > > > > >
> > > > > > * *GitHub Organization*: [https://github.com/opencrawling]
> > > > > > * *Main Code Base*: {{opencrawling/opencrawling}}
> > > > > > * *Specification Repo*: {{opencrawling/open-ingestion-standard}}
> > > > > > * *Documentation & Wiki*: [
> > > > > https://github.com/opencrawling/opencrawling/wiki
> > > > > > ]
> > > > > > * *Java Client SDK*: {{oc-java-client-sdk}} ([Sonatype Central -
> > > > > > org.opencrawling:oc-java-client-sdk|
> > > > > >
> > > > >
> > > >
> > >
> >
> https://central.sonatype.com/artifact/org.opencrawling/oc-java-client-sdk
> > > > > ])
> > > > > > * *Maven Archetypes*: [Sonatype Central -
> > > org.opencrawling.archetypes|
> > > > > >
> > > > >
> > > >
> > >
> >
> https://central.sonatype.com/artifact/org.opencrawling.archetypes/opencrawling-connector-archetypes
> > > > > > ]
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Initial Source & Intellectual Property Submission
> > > > > >
> > > > > > h3. Initial Source Code
> > > > > > The initial codebase to be granted to the ASF resides in the
> > > following
> > > > > > GitHub repositories:
> > > > > > * {{opencrawling/opencrawling}} (Core engine, microservices, UI,
> > > > > > connectors, MCP server)
> > > > > > * {{opencrawling/open-ingestion-standard}} (JSON schemas,
> > whitepaper,
> > > > > > specifications)
> > > > > > * {{opencrawling/opencrawling.github.io}} (Project web site and
> > > > > > documentation source)
> > > > > >
> > > > > > All source code is currently licensed under the *Apache License,
> > > > Version
> > > > > > 2.0*.
> > > > > >
> > > > > > h3. Software Grant / ICLA / CCLA
> > > > > > All core contributors will submit Individual Contributor License
> > > > > Agreements
> > > > > > (ICLAs) and corporate software grants will be executed upon
> > > acceptance
> > > > > into
> > > > > > the Incubator.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. External Dependencies
> > > > > >
> > > > > > All major external dependencies of OpenCrawling use
> > Apache-compatible
> > > > > > licenses (Category A):
> > > > > >
> > > > > > || Dependency || License ||
> > > > > > | *Java Development Kit (JDK 25)* | GPLv2 + Classpath Exception |
> > > > > > | *Spring Boot / Spring AI* | Apache License 2.0 |
> > > > > > | *Apache Tika* | Apache License 2.0 |
> > > > > > | *Apache Kafka Clients* | Apache License 2.0 |
> > > > > > | *Apache Iceberg SDK* | Apache License 2.0 |
> > > > > > | *Apache Ozone Client* | Apache License 2.0 |
> > > > > > | *Jackson / Slf4j / Logback* | Apache 2.0 / MIT / EPL 1.0 |
> > > > > > | *Milvus / Qdrant Java SDKs* | Apache License 2.0 |
> > > > > > | *OpenTelemetry Java SDK* | Apache License 2.0 |
> > > > > > | *React / Vite / Tailwind (Admin UI)* | MIT |
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Cryptography
> > > > > >
> > > > > > OpenCrawling uses standard TLS/HTTPS protocols and hashing
> routines
> > > > > > provided by the standard Java Virtual Machine (JDK) and Spring
> > > Security
> > > > > > framework for secure transport. It does not include custom
> > > > cryptographic
> > > > > > algorithms or controlled export software.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Required Resources
> > > > > >
> > > > > > h3. Mailing Lists
> > > > > > * {{[email protected]}}
> > > > > > * {{[email protected]}}
> > > > > > * {{[email protected]}} (PPMC)
> > > > > >
> > > > > > h3. Git Repositories
> > > > > > * {{https://github.com/apache/incubator-opencrawling}}
> > > > > > * {{https://github.com/apache/incubator-opencrawling-site}}
> > > > > >
> > > > > > h3. Issue Tracking
> > > > > > * GitHub Issues on {{apache/incubator-opencrawling}} (or ASF Jira
> > > > project
> > > > > > {{OPENCRAWLING}})
> > > > > >
> > > > > > h3. CI/CD Infrastructure
> > > > > > * GitHub Actions workflows for automated build, test, multi-arch
> > > Docker
> > > > > > image generation, and Sonar/Scorecard quality checks.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Initial Committers & PPMC Members
> > > > > >
> > > > > > * *Piergiorgio Lucidi* ({{[email protected]}}) – Initial
> > > > Committer
> > > > > &
> > > > > > PPMC
> > > > > > * *Michael Cizmar* ({{[email protected]}}) – Initial
> > > > Committer &
> > > > > > PPMC
> > > > > > * *Luis Cabaceira* ({{[email protected]}}) – Initial
> > > Committer
> > > > &
> > > > > PPMC
> > > > > >
> > > > > > _(Note: Additional mentors and committers will be welcomed during
> > the
> > > > > > discussion period on {{[email protected]}}.)_
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Champions & Mentors
> > > > > >
> > > > > > * *Champion*: Piergiorgio Lucidi ({{[email protected]}}) –
> > ASF
> > > > > Member
> > > > > > * *Mentors*:
> > > > > > ** _(TBD - Interested ASF Members/Incubator PMC members invited
> to
> > > step
> > > > > > forward during proposal discussion)_
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Sponsoring Entity
> > > > > >
> > > > > > The *Apache Incubator PMC* is requested to be the sponsoring
> entity
> > > for
> > > > > > this project.
> > > > >
> > > > >
> ---------------------------------------------------------------------
> > > > > To unsubscribe, e-mail: [email protected]
> > > > > For additional commands, e-mail: [email protected]
> > > > >
> > > > >
> > > > >
> > > >
> > > > Piergiorgio Lucidi
> > > > Mobile: 3395381669
> > > >
> > >
> >
> >
> > --
> > Piergiorgio
> >
>


-- 
Piergiorgio

Reply via email to