This is an automated email from the ASF dual-hosted git repository.

bamaer pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/hop.git


The following commit(s) were added to refs/heads/main by this push:
     new 1413437178 Issue #3478 : Document that copies and partitioning only 
run on the local engine (#8682)
1413437178 is described below

commit 14134371780bca03b5db63b1d03a5ca40919a699
Author: Matt Casters <[email protected]>
AuthorDate: Thu Oct 1 08:04:25 2026 +0200

    Issue #3478 : Document that copies and partitioning only run on the local 
engine (#8682)
---
 .../pages/metadata-types/partition-schema.adoc     |  4 ++++
 .../pipeline/beam/getting-started-with-beam.adoc   |  5 ++++
 .../modules/ROOT/pages/pipeline/partitioning.adoc  | 27 ++++++++++++++++++----
 .../beam-flink-pipeline-engine.adoc                |  4 ++++
 .../modules/ROOT/pages/pipeline/pipelines.adoc     |  4 ++--
 .../spark/getting-started-with-native-spark.adoc   |  2 +-
 .../ROOT/pages/pipeline/specify-copies.adoc        | 21 ++++++++++++++++-
 7 files changed, 59 insertions(+), 8 deletions(-)

diff --git 
a/docs/hop-user-manual/modules/ROOT/pages/metadata-types/partition-schema.adoc 
b/docs/hop-user-manual/modules/ROOT/pages/metadata-types/partition-schema.adoc
index 434eaa0257..fa505829cb 100644
--- 
a/docs/hop-user-manual/modules/ROOT/pages/metadata-types/partition-schema.adoc
+++ 
b/docs/hop-user-manual/modules/ROOT/pages/metadata-types/partition-schema.adoc
@@ -29,6 +29,10 @@ Describes a partition schema.
 A partition schema defines how many ways the row stream will be split.
 The names used for the partitions can be anything you like.
 
+The schema is used only by the local Hop pipeline engine, including when a Hop 
Server runs that engine.
+Beam engines (Flink, Spark, Dataflow and Direct) and the native Spark engine 
ignore it.
+See xref:pipeline/partitioning.adoc#supported-engines[Partitioning].
+
 == Related Plugins
 
 None/All
diff --git 
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/beam/getting-started-with-beam.adoc
 
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/beam/getting-started-with-beam.adoc
index 5dc4e78d9c..a3aed2e4e5 100644
--- 
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/beam/getting-started-with-beam.adoc
+++ 
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/beam/getting-started-with-beam.adoc
@@ -156,6 +156,11 @@ The transforms are in general split into a different types 
described below.
 
 Important to remember is that Beam pipelines try to solve every action in an 
'embarrassingly parallel' way. This means that every transform can and usually 
will run in more than 1 copy.  On large clusters you should expect a lot of 
copies of the same code to run at any given time.
 
+The number from xref:pipeline/specify-copies.adoc[Specify copies] and a 
xref:pipeline/partitioning.adoc[partition schema] do not set that parallelism.
+Both are applied only by the local Hop pipeline engine.
+On Beam, including Flink, a numeric copy count is not a thread count.
+The copies field carries the `SINGLE_BEAM` and `BATCH` flags described below.
+
 === Beam specific transforms
 
 There are a number of Beam specific transforms available which only work on 
the provided Beam pipeline execution engines.
diff --git a/docs/hop-user-manual/modules/ROOT/pages/pipeline/partitioning.adoc 
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/partitioning.adoc
index a58b8da37c..292fac419f 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/pipeline/partitioning.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/pipeline/partitioning.adoc
@@ -15,14 +15,33 @@ specific language governing permissions and limitations
 under the License.
 ////
 :imagesdir: ../assets/images
-:description: Partitioning allows you to distribute all the data from a set 
into distinct subsets according to the rule applied on a table or row, where 
these subsets form a partition of the original set with no item replicated into 
multiple groups.
+:description: Partitioning routes rows to threads on the local Hop pipeline 
engine so related rows stay together. Beam (including Flink) and native Spark 
do not apply a Hop partition schema.
 
 = Partitioning
 
 Partitioning allows you to distribute all the data from a set into distinct 
subsets according to the rule applied on a table or row, where these subsets 
form a partition of the original set with no item replicated into multiple 
groups.
 
-Partitioning data is an important feature to scale your Hop pipelines up and 
out.
-Scaling up makes the most of a single server with multiple CPU cores, while 
scaling out maximizes the resources of multiple servers operating in parallel.
+Partitioning data scales a pipeline on the local Hop pipeline engine.
+It uses the cores of that one engine: each partition is a thread, and rows 
that belong together stay on the same thread.
+It does not spread work across a Flink, Spark or Dataflow cluster.
+
+[[supported-engines]]
+== Supported engines
+
+A partition schema, the copies it starts, repartitioning and swimlanes are 
applied only by the 
xref:pipeline/pipeline-run-configurations/native-local-pipeline-engine.adoc[local
 Hop pipeline engine].
+The 
xref:pipeline/pipeline-run-configurations/native-remote-pipeline-engine.adoc[remote]
 and 
xref:pipeline/pipeline-run-configurations/native-load-balancing-pipeline-engine.adoc[load-balancing]
 engines apply them only when the Hop Server runs the pipeline on that same 
engine.
+The 
xref:pipeline/pipeline-run-configurations/single-threaded-pipeline-engine.adoc[single-threaded]
 engine uses the same routing, but from one thread.
+
+Beam engines do not read the partition schema.
+That includes 
xref:pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc[Flink],
 
xref:pipeline/pipeline-run-configurations/beam-spark-pipeline-engine.adoc[Spark],
 
xref:pipeline/pipeline-run-configurations/beam-dataflow-pipeline-engine.adoc[Dataflow]
 and the 
xref:pipeline/pipeline-run-configurations/beam-direct-pipeline-engine.adoc[Direct]
 runner.
+Setting a schema does not group rows onto the same Flink task or Beam worker.
+Parallelism on those engines comes from the runner.
+The 
xref:pipeline/pipeline-run-configurations/native-spark-pipeline-engine.adoc[native
 Spark] engine ignores the schema as well and uses Spark partitions.
+See xref:pipeline/specify-copies.adoc#supported-engines[Specify copies] for 
the same limit on the copy count.
+
+NOTE: This page is about routing rows to transform copies.
+It is not the table or file split some transforms do from a field value, such 
as xref:pipeline/transforms/tableoutput.adoc[Table output].
+That option belongs to the transform and still runs wherever the transform 
itself is supported.
 
 == Data Partitioning During Processing
 
@@ -33,7 +52,7 @@ The results of which can be verified by examining the preview 
data.
 
 image::hop-gui/pipeline/partitionining-preview.png[Pipeline 
Preview,width="45%"]
 
-To take advantage of the processing resources in your run configuration, you 
can scale up the pipeline using the multi-threading option `Change Number of 
Copies to Start` to produce multiple copies of the transform (click the 
transform's icon to open the context dialog and select 'Specify copies' in the 
'Data routing' category).
+To take advantage of the cores on the local Hop pipeline engine, you can scale 
up the pipeline using the multi-threading option `Change Number of Copies to 
Start` to produce multiple copies of the transform (click the transform's icon 
to open the context dialog and select 'Specify copies' in the 'Data routing' 
category).
 
 As shown below, the x2 notation indicates that two copies will be started at 
runtime.
 
diff --git 
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc
 
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc
index 9462787b24..92b060b191 100644
--- 
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc
+++ 
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc
@@ -26,6 +26,10 @@ This runner allows you to run Hop pipelines on 
https://flink.apache.org[Apache F
 
 The Flink runner supports two modes: Local Direct Flink Runner and Flink 
Runner.
 
+Hop xref:pipeline/specify-copies.adoc[copies] and 
xref:pipeline/partitioning.adoc[partition schemas] are not applied on this 
runner.
+Flink decides how many tasks process each transform.
+Use the parallelism option below, not the copy count or partition schema in 
the pipeline.
+
 The Flink Runner and Flink are suitable for large scale, continuous jobs, and 
provide:
 
 * A streaming-first runtime that supports both batch processing and data 
streaming programs
diff --git a/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipelines.adoc 
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipelines.adoc
index 525d66b592..c5fdc7d765 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipelines.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/pipeline/pipelines.adoc
@@ -79,8 +79,8 @@ Transforms that have to collect a whole set before they can 
emit a row -- xref:p
 
 When a single transform becomes a performance bottleneck (such as a complex 
calculation or slow API call), Hop provides built-in tools to scale:
 
-* **Multiple Copies**: You can tell Hop to run multiple parallel copies of any 
transform (e.g., 4 or 8 copies). Hop will automatically distribute rows across 
these copies to saturate available CPU cores 
(`xref:pipeline/specify-copies.adoc[Specify copies]`).
-* **Data Partitioning**: Partition data streams across distinct worker threads 
or cluster nodes according to partition rules 
(`xref:pipeline/partitioning.adoc[Partitioning]`).
+* **Multiple Copies**: You can tell the local Hop pipeline engine to run 
multiple parallel copies of any transform (e.g., 4 or 8 copies). Hop 
distributes rows across these copies to saturate available CPU cores 
(`xref:pipeline/specify-copies.adoc[Specify copies]`). The copy count is not a 
thread count on Beam (including Flink) or native Spark.
+* **Data Partitioning**: Partition data streams across worker threads of the 
local Hop pipeline engine according to partition rules 
(`xref:pipeline/partitioning.adoc[Partitioning]`). A partition schema is not 
applied on Beam or native Spark.
 * **Row-Level Error Handling**: Route rejected or malformed rows into an 
alternate error stream without stopping the pipeline 
(`xref:pipeline/errorhandling.adoc[Error Handling]`).
 
 == Pipeline Execution Engines
diff --git 
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/spark/getting-started-with-native-spark.adoc
 
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/spark/getting-started-with-native-spark.adoc
index a3a04b898a..3d2a627d2e 100644
--- 
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/spark/getting-started-with-native-spark.adoc
+++ 
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/spark/getting-started-with-native-spark.adoc
@@ -615,7 +615,7 @@ Examples: sorted *Group By* (use *Memory Group By*), 
Beam-only I/O, barriers (*B
 |Reject-to-error style options may not apply; prefer Dataset-oriented dedup 
semantics.
 
 |Transform copies / partitioning GUI
-|Parallelism comes from *Spark partitions*, not Hop “copies” in the local 
sense.
+|Ignored. Parallelism comes from *Spark partitions*, not the local-engine copy 
count or partition schema. See 
xref:pipeline/specify-copies.adoc#supported-engines[Specify copies] and 
xref:pipeline/partitioning.adoc#supported-engines[Partitioning].
 
 |Plugins in mapPartitions
 |Must be present on the Spark classpath (staged list + fat jar). Missing 
plugins fail at runtime on the executor.
diff --git 
a/docs/hop-user-manual/modules/ROOT/pages/pipeline/specify-copies.adoc 
b/docs/hop-user-manual/modules/ROOT/pages/pipeline/specify-copies.adoc
index 03f1f0afc7..6e19f3bcd3 100644
--- a/docs/hop-user-manual/modules/ROOT/pages/pipeline/specify-copies.adoc
+++ b/docs/hop-user-manual/modules/ROOT/pages/pipeline/specify-copies.adoc
@@ -16,7 +16,7 @@ under the License.
 ////
 [[SpecifyCopies]]
 :imagesdir: ../assets/images
-:description: The Specify Copies allows transforms in a pipeline to run with 
multiple copies (threads). This can be used to improve performance when applied 
correctly.
+:description: Specify copies runs a transform in multiple threads on the local 
Hop pipeline engine. Beam (including Flink) and native Spark ignore the copy 
count as a thread count.
 
 = Specify Copies
 
@@ -26,6 +26,25 @@ Having multiple copies of a transform results in multiple 
threads for this trans
 
 WARNING: increasing the number of copies for your transforms is not a silver 
bullet or a `performance=fast` option. Excessive use of the `specify copies` 
option can easily make your pipelines performance worse instead of better.
 
+[[supported-engines]]
+== Supported engines
+
+The copy count starts one thread per copy only on the 
xref:pipeline/pipeline-run-configurations/native-local-pipeline-engine.adoc[local
 Hop pipeline engine].
+The 
xref:pipeline/pipeline-run-configurations/native-remote-pipeline-engine.adoc[remote]
 and 
xref:pipeline/pipeline-run-configurations/native-load-balancing-pipeline-engine.adoc[load-balancing]
 engines apply it only when the pipeline runs with that local engine on a Hop 
Server.
+
+The 
xref:pipeline/pipeline-run-configurations/single-threaded-pipeline-engine.adoc[single-threaded]
 engine creates the same copies, but it drives every transform from one thread, 
so the copies do not run in parallel.
+
+Beam engines do not use the number as a thread count.
+That includes 
xref:pipeline/pipeline-run-configurations/beam-flink-pipeline-engine.adoc[Flink],
 
xref:pipeline/pipeline-run-configurations/beam-spark-pipeline-engine.adoc[Spark],
 
xref:pipeline/pipeline-run-configurations/beam-dataflow-pipeline-engine.adoc[Dataflow]
 and the 
xref:pipeline/pipeline-run-configurations/beam-direct-pipeline-engine.adoc[Direct]
 runner.
+Those runners choose their own parallelism.
+On Beam the copies field is a flag string, not a thread count: `SINGLE_BEAM` 
asks the engine to try to run the transform once (Flink still runs that work in 
parallel), and `BATCH` batches rows before they reach the transform.
+See xref:pipeline/beam/getting-started-with-beam.adoc[Getting started with 
Apache Beam].
+xref:pipeline/transforms/rowgenerator.adoc[Generate rows] is the one transform 
that reads a numeric copy count on Beam, and only as the number of initial 
bundles of its synthetic source.
+
+The 
xref:pipeline/pipeline-run-configurations/native-spark-pipeline-engine.adoc[native
 Spark] engine ignores the copy count.
+Its parallelism comes from Spark partitions.
+See xref:pipeline/spark/getting-started-with-native-spark.adoc[Getting started 
with the native Spark pipeline engine].
+
 == Changing the number of copies for a transform
 
 Click on a transform's icon and click on the `Specify copies` icon in the 
pop-up dialog.

Reply via email to