This is an automated email from the ASF dual-hosted git repository.
bamaer pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/hop.git
The following commit(s) were added to refs/heads/main by this push:
new ea38e959c4 Issue #2408 : Add an Apache Hop sizing guide (#8675)
ea38e959c4 is described below
commit ea38e959c42fde603ae31518584060094461843d
Author: Matt Casters <[email protected]>
AuthorDate: Thu Oct 1 08:55:05 2026 +0200
Issue #2408 : Add an Apache Hop sizing guide (#8675)
* Issue #2408 : Add an Apache Hop sizing guide
* Fixes #2408 : Fix rowset variable and Hop Server log retention defaults
in sizing guide
---------
Co-authored-by: Bart Maertens <[email protected]>
---
docs/hop-user-manual/modules/ROOT/nav.adoc | 1 +
.../modules/ROOT/pages/docker-container.adoc | 2 +-
.../ROOT/pages/installation-configuration.adoc | 4 +-
.../hop-user-manual/modules/ROOT/pages/sizing.adoc | 259 +++++++++++++++++++++
4 files changed, 264 insertions(+), 2 deletions(-)
diff --git a/docs/hop-user-manual/modules/ROOT/nav.adoc
b/docs/hop-user-manual/modules/ROOT/nav.adoc
index 7d03e93591..63672024f5 100644
--- a/docs/hop-user-manual/modules/ROOT/nav.adoc
+++ b/docs/hop-user-manual/modules/ROOT/nav.adoc
@@ -30,6 +30,7 @@ under the License.
** xref:hop-vs-kettle/import-kettle-projects.adoc[Upgrade Kettle to Hop]
* xref:concepts.adoc[Concepts]
* xref:installation-configuration.adoc[Installation and Configuration]
+* xref:sizing.adoc[Sizing guide]
* xref:docker-container.adoc[Hop in Docker]
** xref:hop-web-docker.adoc[Hop Web in Docker]
* xref:cloud.adoc[Hop in the Cloud]
diff --git a/docs/hop-user-manual/modules/ROOT/pages/docker-container.adoc
b/docs/hop-user-manual/modules/ROOT/pages/docker-container.adoc
index a9d5bca828..25ca98a996 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/docker-container.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/docker-container.adoc
@@ -246,7 +246,7 @@ Below are the variables you can use for a **long-lived**
container, running Hop
|The maximum number of log lines kept in memory by the server.
|```HOP_SERVER_MAX_LOG_TIMEOUT```
-|`0` (never clean up log lines)
+|`1440` (clean up log lines after one day)
|The time (in minutes) it takes for a log line to be cleaned up in memory.
|```HOP_SERVER_MAX_OBJECT_TIMEOUT```
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/installation-configuration.adoc
b/docs/hop-user-manual/modules/ROOT/pages/installation-configuration.adoc
index 37f6549441..26e50b29d5 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/installation-configuration.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/installation-configuration.adoc
@@ -36,6 +36,8 @@ Hop's limited footprint should allow it to run on any modern
physical or virtual
For the default Hop distribution, a minimum of 1 CPU/core and 4GB RAM should
do, even though you can tweak Hop to run on machines with even less memory.
+xref:sizing.adoc[Sizing Apache Hop] splits that floor into three cases: a
stripped `hop-run` on a small device, a workstation for Hop Gui, and batch runs
on Spark or Beam.
+
Hop Runs on the following operating systems:
* Windows 7 or higher
@@ -137,7 +139,7 @@ TIP: For Hop Web in Docker, persist `HOP_CONFIG_FOLDER` and
project homes on vol
== Additional configuration
-=== JVM memory settings
+=== JVM memory settings [[JvmMemorySettings]]
By default, Hop only sets a maximum for the JVM Heap size Hop can allocate.
diff --git a/docs/hop-user-manual/modules/ROOT/pages/sizing.adoc
b/docs/hop-user-manual/modules/ROOT/pages/sizing.adoc
new file mode 100644
index 0000000000..f417f267ac
--- /dev/null
+++ b/docs/hop-user-manual/modules/ROOT/pages/sizing.adoc
@@ -0,0 +1,259 @@
+////
+Licensed to the Apache Software Foundation (ASF) under one
+or more contributor license agreements. See the NOTICE file
+distributed with this work for additional information
+regarding copyright ownership. The ASF licenses this file
+to you under the Apache License, Version 2.0 (the
+"License"); you may not use this file except in compliance
+with the License. You may obtain a copy of the License at
+ http://www.apache.org/licenses/LICENSE-2.0
+Unless required by applicable law or agreed to in writing,
+software distributed under the License is distributed on an
+"AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+KIND, either express or implied. See the License for the
+specific language governing permissions and limitations
+under the License.
+////
+[[SizingGuide]]
+:description: How to size Apache Hop for a small device, a data engineer
workstation, and a big-data engine such as Spark.
+
+= Sizing Apache Hop
+
+Hop does not reserve a fixed amount of memory or disk for a pipeline.
+What you need depends on which tool you start, which plugins you keep, and
whether a transform holds rows or passes them through.
+
+xref:installation-configuration.adoc[Installation and configuration] is the
floor for a default client: Java 21, one CPU core and 4 GB of RAM.
+This page is the next step.
+It covers a stripped `hop-run` on a small device, a workstation for day-to-day
design, and batch pipelines that run on Spark or Beam.
+
+== What uses memory and disk
+
+The launch scripts set a maximum heap of 2 GB when `HOP_OPTIONS` is empty
(`-Xmx2048m` in `hop-gui`, `hop-run` and `hop-server`).
+That flag is a ceiling.
+The JVM does not allocate 2 GB at startup.
+Metaspace, thread stacks and mapped jars sit outside the heap, so the process
is always larger than the heap you set.
+Set `HOP_OPTIONS` in the environment when you want a different ceiling.
+That value replaces the default in the scripts.
+xref:installation-configuration.adoc#JvmMemorySettings[JVM memory settings]
shows the values the scripts accept.
+
+On the local engine, rows stream from one transform to the next.
+Each hop buffers up to the row set size, ten thousand rows by default (`Row
set size` on the
xref:pipeline/pipeline-run-configurations/native-local-pipeline-engine.adoc[local
pipeline engine]).
+A streaming pipeline's heap follows the width of one row times that buffer,
not the size of the source, unless a transform accumulates.
+
+These transforms do keep data:
+
+* xref:pipeline/transforms/memgroupby.adoc[Memory Group By] keeps every group
in memory.
+* xref:pipeline/transforms/sort.adoc[Sort rows] keeps a batch in memory and
then spills to the sort directory.
+ Leave free disk there for large sorts.
+* xref:pipeline/transforms/databaselookup.adoc[Database lookup] with a cache,
and especially *Load all data from table*, keeps lookup rows in memory.
+ A cache size of 0 means the whole table.
+* xref:pipeline/transforms/streamlookup.adoc[Stream lookup] keeps the lookup
stream in memory.
+* xref:pipeline/transforms/jsoninput.adoc[JSON input] parses each file into a
document before it emits rows.
+ Size the heap for the largest file, not for one output row.
+* xref:pipeline/transforms/excelinput.adoc[Excel input] loads a workbook when
you use the default spreadsheet type.
+ *Excel XLSX (Streaming)* is the low-memory choice for large `.xlsx` files.
+* xref:pipeline/transforms/rest.adoc[REST client] holds the current response
body in a field.
+ One request is sent per input row.
+
+The default `local` run configuration also records execution information with
the `first-last` data profile (the first and last 100 rows of each transform).
+xref:hop-server/index.adoc[Hop Server] keeps every log line in memory for one
day by default (`max_log_lines` is 0, `max_log_timeout_minutes` is 1440).
+Set `max_log_lines` in the server configuration to cap it, see
xref:hop-server/index.adoc#_cleanup_settings_and_where_they_come_from[Cleanup
settings].
+In the xref:docker-container.adoc[Docker] image, `HOP_SERVER_MAX_LOG_LINES`
and `HOP_SERVER_MAX_LOG_TIMEOUT` set the same two values.
+
+Disk for the client itself is the unzipped tree: `lib/core`, `lib/jdbc`,
`lib/swt` and `plugins/`.
+`lib/core` is on the classpath of every tool.
+Do not delete individual jars from it.
+`lib/jdbc` is the shared JDBC folder (about 30 MB in a current client) and is
only needed for databases you actually use.
+`lib/swt` ships a copy for each operating system (about 15 MB together, a few
MB for one).
+`hop-run` puts the current operating system's SWT jars on the classpath and
does not open a window.
+Every other folder under `plugins/` is optional.
+Removing a plugin that a pipeline or workflow still references makes that run
fail at startup.
+See xref:best-practices/index.adoc[Best practices] (*Remove unused plugins*).
+
+== Edge nodes
+
+Use `hop-run` for a device that only executes a pipeline.
+You do not need Hop Gui, a display, or the plugins you are not calling.
+
+=== Worked example
+
+The example pipeline has two transforms.
+xref:pipeline/transforms/jsoninput.adoc[JSON input] reads a file of objects.
+xref:pipeline/transforms/rest.adoc[REST client] POSTs one field from each row
to an HTTP endpoint.
+It was run with the default `local` run configuration (`hop-run -j default -r
local`) on Linux x86_64 and Java 21.
+
+These plugins had to stay:
+
+* `plugins/transforms/json`
+* `plugins/transforms/rest`
+* `plugins/misc/rest` (REST client settings, shared with the REST transform)
+* `plugins/misc/projects` (so `-j` / `--project` can load the `local` run
configuration)
+
+Every other folder under `plugins/` was removed.
+`lib/core` and the Linux SWT jars stayed.
+`lib/jdbc` was left out because the pipeline uses no database.
+
+Repeat the check on the release you deploy.
+Plugin jars move between versions.
+`du -sm` on the install tree and `hop-run` with a small file will tell you if
the floor moved.
+
+=== Measured footprint
+
+[options="header",cols="2,1,1"]
+|===
+|Layout |Disk |Notes
+
+|Client archive
+|about 365 MB
+|Compressed zip of the client.
+
+|Client, unpacked
+|about 440 MB
+|`lib/` about 240 MB, `plugins/` about 190 MB.
+
+|`lib/core`
+|about 200 MB
+|Shared classpath. Keep it.
+
+|`plugins/transforms/script`
+|about 120 MB
+|Largest single plugin folder in that client. Safe to remove when you do not
script.
+
+|Stripped JSON-to-REST layout
+|about 215 MB
+|`lib/core`, Linux SWT, and the four plugin folders above.
+|===
+
+Resident set of that `hop-run` (RSS, not just the heap):
+
+[options="header",cols="2,1,1,1"]
+|===
+|Run |Smallest heap that finished |Heap that failed |Resident set
+
+|20 small JSON objects
+|`-Xmx32m`
+|`-Xmx24m`
+|about 200 MB
+
+|20,000 objects, 4.6 MB file
+|`-Xmx48m`
+|not measured below 48 MB
+|about 270 MB
+
+|Same 20-row pipeline, all client plugins still installed
+|`-Xmx64m`
+|`-Xmx32m`
+|about 290 MB at 64 MB heap, about 400 MB at 512 MB heap
+|===
+
+The 4.6 MB file still fitted in a 48 MB heap because the rows were small and
were sent on immediately.
+A wide or deeply nested document is larger in memory than it is on disk.
+Raise the heap for the biggest file you will parse, then leave a margin.
+
+=== What to plan for
+
+[options="header",cols="1,2"]
+|===
+|Resource |Plan
+
+|Disk
+|512 MB free on the device.
+The stripped tree was about 215 MB.
+The rest is the JSON you read, the audit folder, and a later plugin you did
not think of yet.
+
+|Memory
+|512 MB of RAM for the process and the operating system.
+Set `HOP_OPTIONS` to `-Xmx128m`.
+32 MB of heap was enough for the tiny file and is too tight for a real
document.
+
+|CPU
+|One core.
+The local engine uses one thread per transform copy, and this pipeline has two
transforms.
+
+|Display
+|None.
+`hop-run` does not open a window.
+
+|Network
+|Enough bandwidth for the REST calls.
+This pipeline sends one request per row, so latency dominates once the files
are small.
+|===
+
+== A data engineer workstation
+
+Designing pipelines in Hop Gui is a different budget from the edge layout.
+The GUI, the full plugin set and a 2 GB heap are the normal case.
+The xref:installation-configuration.adoc[installation] floor (one core, 4 GB
of RAM) still runs that client.
+A machine you work on all day should sit above the floor.
+
+[options="header",cols="1,2"]
+|===
+|Resource |Plan
+
+|Memory
+|8 GB of RAM.
+That fits the default 2 GB heap, metaspace, the operating system, and a
browser or a database tool.
+Raise `HOP_OPTIONS` (for example `-Xmx4g`) when a sort, a lookup cache, a JSON
file or an Excel sheet runs out of heap.
+Do not raise it to hide a transform that is holding the whole data set.
+Stream, or spill to disk, instead.
+
+|CPU
+|4 cores.
+Each transform copy is a thread, and the GUI needs time while a pipeline is
running.
+One core is the floor and will feel stuck as soon as several transforms are
busy.
+
+|Disk
+|A few gigabytes free.
+The unpacked client is under 1 GB.
+Keep room for a second Hop version side by side, JDBC drivers, project files,
logs and sort spills.
+An SSD matters once sorts spill.
+
+|Network
+|Low latency to the databases you query row by row.
+xref:pipeline/transforms/databaselookup.adoc[Database lookup] and
xref:pipeline/transforms/databasejoin.adoc[Database join] wait on a round trip
per row (or per cached miss).
+Bulk writes are limited by bandwidth instead.
+Run Hop close to those databases when the pipeline is chatty.
+
+|Display
+|Full HD, 1920×1080.
+Hop Gui is a desktop canvas: the graph, the properties and the log are meant
to be on screen together.
+The same layout is what you look at in xref:hop-gui/hop-web.adoc[Hop Web].
+|===
+
+== Big data
+
+A pipeline that is larger than the workstation should not be forced through
the local engine by raising `HOP_OPTIONS`.
+Pick a run configuration that hands the data to another engine:
+
+* xref:pipeline/spark/getting-started-with-native-spark.adoc[Native Spark] for
batch pipelines on Spark 4.1.
+ The cluster JVM must be able to run Java 21, the same as Hop.
+* xref:pipeline/beam/getting-started-with-beam.adoc[Apache Beam] when you need
Spark 3.5, Flink, Dataflow, streaming or windows.
+
+The native Spark and Beam engines are optional plugins.
+They are not in the client sizes in the table above.
+The native Spark engine plugin is about 340 MB on disk.
+The Beam engine plugin archive is about 390 MB compressed, before you unpack
it.
+After you install one, Hop Gui is still the workstation in the previous
section.
+It builds and submits the job.
+The dataset stays in the cluster.
+
+On native Spark, a transform is either a Spark operation or a small Hop
pipeline run once per partition (`mapPartitions`).
+Size Spark executor memory for one partition plus shuffle, not by giving the
Hop process a larger heap.
+A wrapped transform that would buffer the whole data set on the local engine
only sees the partition it was given.
+Aggregations that have a native implementation (sort, merge join, memory group
by, and the Spark file and lake transforms) run as Spark operations across the
data set.
+The run-configuration templates for a standalone master and for YARN fill in 2
GB for the driver, 2 GB for each executor and 2 executor cores as a starting
point, not as a requirement.
+Change those to match the cluster.
+Streaming and windowing stay on Beam.
+
+== Checking an installation
+
+Measure the tree you actually ship:
+
+[source,shell]
+----
+du -sm hop hop/lib hop/lib/core hop/lib/jdbc hop/lib/swt hop/plugins
+----
+
+Run the smallest pipeline you care about with a low `HOP_OPTIONS` and raise
the heap only until it finishes.
+`Java heap space` in the log is the heap ceiling.
+A process that is killed by the operating system while the heap still had room
is metaspace or native memory: the resident set, not `-Xmx`, is what the device
must provide.