This is an automated email from the ASF dual-hosted git repository.
bamaer pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/hop.git
The following commit(s) were added to refs/heads/main by this push:
new 1413437178 Issue #3478 : Document that copies and partitioning only
run on the local engine (#8682)
1413437178 is described below
commit 14134371780bca03b5db63b1d03a5ca40919a699
Author: Matt Casters <[email protected]>
AuthorDate: Thu Oct 1 08:04:25 2026 +0200
Issue #3478 : Document that copies and partitioning only run on the local
engine (#8682)
---
.../pages/metadata-types/partition-schema.adoc | 4 ++++
.../pipeline/beam/getting-started-with-beam.adoc | 5 ++++
.../modules/ROOT/pages/pipeline/partitioning.adoc | 27 ++++++++++++++++++----
.../beam-flink-pipeline-engine.adoc | 4 ++++
.../modules/ROOT/pages/pipeline/pipelines.adoc | 4 ++--
.../spark/getting-started-with-native-spark.adoc | 2 +-
.../ROOT/pages/pipeline/specify-copies.adoc | 21 ++++++++++++++++-
7 files changed, 59 insertions(+), 8 deletions(-)
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/metadata-types/partition-schema.adoc
b/docs/hop-user-manual/modules/ROOT/pages/metadata-types/partition-schema.adoc
index 434eaa0257..fa505829cb 100644
---
a/docs/hop-user-manual/modules/ROOT/pages/metadata-types/partition-schema.adoc
+++
b/docs/hop-user-manual/modules/ROOT/pages/metadata-types/partition-schema.adoc
@@ -29,6 +29,10 @@ Describes a partition schema.
A partition schema defines how many ways the row stream will be split.
The names used for the partitions can be anything you like.
+The schema is used only by the local Hop pipeline engine, including when a Hop
Server runs that engine.
+Beam engines (Flink, Spark, Dataflow and Direct) and the native Spark engine
ignore it.
+See xref:pipeline/partitioning.adoc#supported-engines[Partitioning].
+
== Related Plugins
None/All
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/beam/getting-started-with-beam.adoc
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/beam/getting-started-with-beam.adoc
index 5dc4e78d9c..a3aed2e4e5 100644
---
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/beam/getting-started-with-beam.adoc
+++
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/beam/getting-started-with-beam.adoc
@@ -156,6 +156,11 @@ The transforms are in general split into a different types
described below.
Important to remember is that Beam pipelines try to solve every action in an
'embarrassingly parallel' way. This means that every transform can and usually
will run in more than 1 copy. On large clusters you should expect a lot of
copies of the same code to run at any given time.
+The number from xref:pipeline/specify-copies.adoc[Specify copies] and a
xref:pipeline/partitioning.adoc[partition schema] do not set that parallelism.
+Both are applied only by the local Hop pipeline engine.
+On Beam, including Flink, a numeric copy count is not a thread count.
+The copies field carries the `SINGLE_BEAM` and `BATCH` flags described below.
+
=== Beam specific transforms
There are a number of Beam specific transforms available which only work on
the provided Beam pipeline execution engines.
diff --git a/docs/hop-user-manual/modules/ROOT/pages/pipeline/partitioning.adoc
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/partitioning.adoc
index a58b8da37c..292fac419f 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/pipeline/partitioning.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/pipeline/partitioning.adoc
@@ -15,14 +15,33 @@ specific language governing permissions and limitations
under the License.
////
:imagesdir: ../assets/images
-:description: Partitioning allows you to distribute all the data from a set
into distinct subsets according to the rule applied on a table or row, where
these subsets form a partition of the original set with no item replicated into
multiple groups.
+:description: Partitioning routes rows to threads on the local Hop pipeline
engine so related rows stay together. Beam (including Flink) and native Spark
do not apply a Hop partition schema.
= Partitioning
Partitioning allows you to distribute all the data from a set into distinct
subsets according to the rule applied on a table or row, where these subsets
form a partition of the original set with no item replicated into multiple
groups.
-Partitioning data is an important feature to scale your Hop pipelines up and
out.
-Scaling up makes the most of a single server with multiple CPU cores, while
scaling out maximizes the resources of multiple servers operating in parallel.
+Partitioning data scales a pipeline on the local Hop pipeline engine.
+It uses the cores of that one engine: each partition is a thread, and rows
that belong together stay on the same thread.
+It does not spread work across a Flink, Spark or Dataflow cluster.
+
+[[supported-engines]]
+== Supported engines
+
+A partition schema, the copies it starts, repartitioning and swimlanes are
applied only by the
xref:pipeline/pipeline-run-configurations/native-local-pipeline-engine.adoc[local
Hop pipeline engine].
+The
xref:pipeline/pipeline-run-configurations/native-remote-pipeline-engine.adoc[remote]
and
xref:pipeline/pipeline-run-configurations/native-load-balancing-pipeline-engine.adoc[load-balancing]
engines apply them only when the Hop Server runs the pipeline on that same
engine.
+The
xref:pipeline/pipeline-run-configurations/single-threaded-pipeline-engine.adoc[single-threaded]
engine uses the same routing, but from one thread.
+
+Beam engines do not read the partition schema.
+That includes
xref:pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc[Flink],
xref:pipeline/pipeline-run-configurations/beam-spark-pipeline-engine.adoc[Spark],
xref:pipeline/pipeline-run-configurations/beam-dataflow-pipeline-engine.adoc[Dataflow]
and the
xref:pipeline/pipeline-run-configurations/beam-direct-pipeline-engine.adoc[Direct]
runner.
+Setting a schema does not group rows onto the same Flink task or Beam worker.
+Parallelism on those engines comes from the runner.
+The
xref:pipeline/pipeline-run-configurations/native-spark-pipeline-engine.adoc[native
Spark] engine ignores the schema as well and uses Spark partitions.
+See xref:pipeline/specify-copies.adoc#supported-engines[Specify copies] for
the same limit on the copy count.
+
+NOTE: This page is about routing rows to transform copies.
+It is not the table or file split some transforms do from a field value, such
as xref:pipeline/transforms/tableoutput.adoc[Table output].
+That option belongs to the transform and still runs wherever the transform
itself is supported.
== Data Partitioning During Processing
@@ -33,7 +52,7 @@ The results of which can be verified by examining the preview
data.
image::hop-gui/pipeline/partitionining-preview.png[Pipeline
Preview,width="45%"]
-To take advantage of the processing resources in your run configuration, you
can scale up the pipeline using the multi-threading option `Change Number of
Copies to Start` to produce multiple copies of the transform (click the
transform's icon to open the context dialog and select 'Specify copies' in the
'Data routing' category).
+To take advantage of the cores on the local Hop pipeline engine, you can scale
up the pipeline using the multi-threading option `Change Number of Copies to
Start` to produce multiple copies of the transform (click the transform's icon
to open the context dialog and select 'Specify copies' in the 'Data routing'
category).
As shown below, the x2 notation indicates that two copies will be started at
runtime.
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc
index 9462787b24..92b060b191 100644
---
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc
+++
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc
@@ -26,6 +26,10 @@ This runner allows you to run Hop pipelines on
https://flink.apache.org[Apache F
The Flink runner supports two modes: Local Direct Flink Runner and Flink
Runner.
+Hop xref:pipeline/specify-copies.adoc[copies] and
xref:pipeline/partitioning.adoc[partition schemas] are not applied on this
runner.
+Flink decides how many tasks process each transform.
+Use the parallelism option below, not the copy count or partition schema in
the pipeline.
+
The Flink Runner and Flink are suitable for large scale, continuous jobs, and
provide:
* A streaming-first runtime that supports both batch processing and data
streaming programs
diff --git a/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipelines.adoc
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipelines.adoc
index 525d66b592..c5fdc7d765 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipelines.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipelines.adoc
@@ -79,8 +79,8 @@ Transforms that have to collect a whole set before they can
emit a row -- xref:p
When a single transform becomes a performance bottleneck (such as a complex
calculation or slow API call), Hop provides built-in tools to scale:
-* **Multiple Copies**: You can tell Hop to run multiple parallel copies of any
transform (e.g., 4 or 8 copies). Hop will automatically distribute rows across
these copies to saturate available CPU cores
(`xref:pipeline/specify-copies.adoc[Specify copies]`).
-* **Data Partitioning**: Partition data streams across distinct worker threads
or cluster nodes according to partition rules
(`xref:pipeline/partitioning.adoc[Partitioning]`).
+* **Multiple Copies**: You can tell the local Hop pipeline engine to run
multiple parallel copies of any transform (e.g., 4 or 8 copies). Hop
distributes rows across these copies to saturate available CPU cores
(`xref:pipeline/specify-copies.adoc[Specify copies]`). The copy count is not a
thread count on Beam (including Flink) or native Spark.
+* **Data Partitioning**: Partition data streams across worker threads of the
local Hop pipeline engine according to partition rules
(`xref:pipeline/partitioning.adoc[Partitioning]`). A partition schema is not
applied on Beam or native Spark.
* **Row-Level Error Handling**: Route rejected or malformed rows into an
alternate error stream without stopping the pipeline
(`xref:pipeline/errorhandling.adoc[Error Handling]`).
== Pipeline Execution Engines
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/spark/getting-started-with-native-spark.adoc
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/spark/getting-started-with-native-spark.adoc
index a3a04b898a..3d2a627d2e 100644
---
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/spark/getting-started-with-native-spark.adoc
+++
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/spark/getting-started-with-native-spark.adoc
@@ -615,7 +615,7 @@ Examples: sorted *Group By* (use *Memory Group By*),
Beam-only I/O, barriers (*B
|Reject-to-error style options may not apply; prefer Dataset-oriented dedup
semantics.
|Transform copies / partitioning GUI
-|Parallelism comes from *Spark partitions*, not Hop “copies” in the local
sense.
+|Ignored. Parallelism comes from *Spark partitions*, not the local-engine copy
count or partition schema. See
xref:pipeline/specify-copies.adoc#supported-engines[Specify copies] and
xref:pipeline/partitioning.adoc#supported-engines[Partitioning].
|Plugins in mapPartitions
|Must be present on the Spark classpath (staged list + fat jar). Missing
plugins fail at runtime on the executor.
diff --git
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/specify-copies.adoc
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/specify-copies.adoc
index 03f1f0afc7..6e19f3bcd3 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/pipeline/specify-copies.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/pipeline/specify-copies.adoc
@@ -16,7 +16,7 @@ under the License.
////
[[SpecifyCopies]]
:imagesdir: ../assets/images
-:description: The Specify Copies allows transforms in a pipeline to run with
multiple copies (threads). This can be used to improve performance when applied
correctly.
+:description: Specify copies runs a transform in multiple threads on the local
Hop pipeline engine. Beam (including Flink) and native Spark ignore the copy
count as a thread count.
= Specify Copies
@@ -26,6 +26,25 @@ Having multiple copies of a transform results in multiple
threads for this trans
WARNING: increasing the number of copies for your transforms is not a silver
bullet or a `performance=fast` option. Excessive use of the `specify copies`
option can easily make your pipelines performance worse instead of better.
+[[supported-engines]]
+== Supported engines
+
+The copy count starts one thread per copy only on the
xref:pipeline/pipeline-run-configurations/native-local-pipeline-engine.adoc[local
Hop pipeline engine].
+The
xref:pipeline/pipeline-run-configurations/native-remote-pipeline-engine.adoc[remote]
and
xref:pipeline/pipeline-run-configurations/native-load-balancing-pipeline-engine.adoc[load-balancing]
engines apply it only when the pipeline runs with that local engine on a Hop
Server.
+
+The
xref:pipeline/pipeline-run-configurations/single-threaded-pipeline-engine.adoc[single-threaded]
engine creates the same copies, but it drives every transform from one thread,
so the copies do not run in parallel.
+
+Beam engines do not use the number as a thread count.
+That includes
xref:pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc[Flink],
xref:pipeline/pipeline-run-configurations/beam-spark-pipeline-engine.adoc[Spark],
xref:pipeline/pipeline-run-configurations/beam-dataflow-pipeline-engine.adoc[Dataflow]
and the
xref:pipeline/pipeline-run-configurations/beam-direct-pipeline-engine.adoc[Direct]
runner.
+Those runners choose their own parallelism.
+On Beam the copies field is a flag string, not a thread count: `SINGLE_BEAM`
asks the engine to try to run the transform once (Flink still runs that work in
parallel), and `BATCH` batches rows before they reach the transform.
+See xref:pipeline/beam/getting-started-with-beam.adoc[Getting started with
Apache Beam].
+xref:pipeline/transforms/rowgenerator.adoc[Generate rows] is the one transform
that reads a numeric copy count on Beam, and only as the number of initial
bundles of its synthetic source.
+
+The
xref:pipeline/pipeline-run-configurations/native-spark-pipeline-engine.adoc[native
Spark] engine ignores the copy count.
+Its parallelism comes from Spark partitions.
+See xref:pipeline/spark/getting-started-with-native-spark.adoc[Getting started
with the native Spark pipeline engine].
+
== Changing the number of copies for a transform
Click on a transform's icon and click on the `Specify copies` icon in the
pop-up dialog.