Hi all, Flume 2.0.0 is getting close, and before I set up the release process I'd appreciate your advice on how to structure it. Many of you have far more experience than me with releasing multi-repository projects like Log4j, so feel free to point out anything that looks off.
Flume is now organized like Log4j: - `logging-flume`: the core components, including the executable Flume Node. It currently also hosts `flume-parent`, the BOMs and the (disabled) binary distribution. - 14 satellite repositories with the sources, sinks and channels that carry heavy dependencies (Avro, Thrift, Hadoop, Kafka, Spring Boot, ...). I expect the core to change rarely, while the satellites will need releases that follow the lifecycle of their dependencies. So I'd like to release the satellites independently, without waiting for (or forcing) a core release. The catch is that the core repository still points at the satellites: its BOM lists their artifacts, and the distribution bundles them, so a complete, consistent core release would have to come after every round of satellite releases. ## What I'm planning My idea is to make dependencies point one way only: satellites depend on the core, never the reverse. Everything that aggregates the satellites would move to a new repository, released last. 1. `logging-flume` -> **Apache Flume Core** (ATR key `logging-flume-core`) - Core modules only. The non-trivial parts of `flume-parent` (Surefire, PMD, Spotless configuration) would move to `logging-parent`, and the dependencies only some satellites need for tests (e.g., the Hadoop mini-cluster) would move to those satellites. What remains is little more than metadata, so `flume-parent` can be dropped and all repositories can inherit from `logging-parent` directly. - A `flume-core-bom` listing only the artifacts this repository builds. - No more references to satellite artifacts. 2. Satellite repositories -> one subproject each. They depend on a released core through `flume-core-bom` and have their own release cadence. 3. New `logging-flume-dist` -> **Apache Flume**, released whenever there is a tested set worth publishing: - `flume-bom`, listing compatible versions of the core and all satellites; - a binary distribution recipe that users can fork and extend with their own plugins; - a Docker image definition. None of the BOMs has been released yet, so they can still be renamed or split. The release order for 2.0.0 would be: core, then satellites, then dist. ## Binary distribution and Docker The 1.x ZIP bundled every module, yet users usually had to add their own plugins anyway. So I'm leaning towards publishing a recipe (an assembly project to copy) and a Docker image instead of a monolithic ZIP. For Docker, my reading of the release policy is that convenience binaries must only contain artifacts built from the released sources. Rebuilding the image on an updated base image doesn't change any Flume bits, so I think it can be rebuilt on a schedule from CI without a new vote, with a vote only when the recipe or the Flume version changes. I'm not sure about this, though. ## Where I'd like your advice 1. Does this split make sense to you, or would you organize it differently? In particular, are you fine with moving the Flume build configuration to `logging-parent`? 2. Should the satellites keep the core version number (2.0.x) or get independent version numbers? I would rather version them separately. 3. Is a binary ZIP still worth producing, or are a recipe and a Docker image enough? Thanks! Piotr
