Abacn commented on code in PR #40298:
URL: https://github.com/apache/beam/pull/40298#discussion_r4124378894


##########
runners/prism/java/src/main/java/org/apache/beam/runners/prism/PrismRunner.java:
##########
@@ -82,19 +83,37 @@ public PipelineResult run(Pipeline pipeline) {
         prismPipelineOptions.getDefaultEnvironmentType(),
         prismPipelineOptions.getJobEndpoint());
 
-    try {
-      PrismExecutor executor = startPrism();
-      PortableRunner delegate = 
PortableRunner.fromOptions(prismPipelineOptions);
-      return new PrismPipelineResult(delegate.run(pipeline), executor::stop);
-    } catch (IOException e) {
-      throw new RuntimeException(e);
+    for (int attempt = 1; ; attempt++) {
+      PrismExecutor executor;
+      try {
+        executor = startPrism();
+      } catch (IOException e) {
+        throw new RuntimeException(e);
+      }
+      try {
+        PortableRunner delegate = 
PortableRunner.fromOptions(prismPipelineOptions);
+        return new PrismPipelineResult(delegate.run(pipeline), executor::stop);
+      } catch (RuntimeException e) {
+        boolean alive = executor.isAlive();
+        executor.stop();
+        if (attempt < MAX_START_ATTEMPTS
+            && (!alive
+                || (e.getMessage() != null && 
e.getMessage().contains("JobService/Prepare")))) {

Review Comment:
   Retry can reduce flakiness in general. However on genuine failure, keep 
retrying only delays the program execution. We should only retry transient 
failures (e.g. port conflict)



##########
runners/spark/src/main/java/org/apache/beam/runners/spark/structuredstreaming/translation/SparkSessionFactory.java:
##########
@@ -176,19 +179,22 @@ private static SparkSession.Builder sessionBuilder(
         sparkConf.setAppName(options.getAppName());
       }
 
-      if (options.getFilesToStage() != null && 
!options.getFilesToStage().isEmpty()) {
-        // Append the files to stage provided by the user to `spark.jars`.
-        PipelineResources.prepareFilesForStaging(options);
-        String[] filesToStage = filterFilesToStage(options, 
Collections.emptyList());
-        String[] jars = getSparkJars(sparkConf);
-        sparkConf.setJars(jars.length > 0 ? ArrayUtils.addAll(jars, 
filesToStage) : filesToStage);
-      } else if (!sparkConf.contains("spark.jars") && 
!master.startsWith("local[")) {
-        // Stage classpath if `spark.jars` not set and not in local mode.
-        PipelineResources.prepareFilesForStaging(options);
-        // Set `spark.jars`, exclude JRE libs and jars causing conflicts using 
`userClassPathFirst`.
-        sparkConf.setJars(filterFilesToStage(options, SPARK_JAR_EXCLUDES));
-        // Enable `userClassPathFirst` to prevent issues with guava, jackson 
and others.
-        sparkConf.setIfMissing("spark.executor.userClassPathFirst", "true");
+      if (!master.startsWith("local")) {

Review Comment:
   Some flakiness should be fixed by #40270 already



##########
runners/spark/src/main/java/org/apache/beam/runners/spark/structuredstreaming/SparkStructuredStreamingRunner.java:
##########
@@ -177,8 +177,13 @@ public SparkStructuredStreamingPipelineResult run(final 
Pipeline pipeline) {
                 ctxRef.set(ctx);
                 if (!cancelRequested.get()) {
                   ctx.evaluate();
+                } else {
+                  cancelSparkJobs.run();
                 }
               } finally {
+                if (cancelRequested.get()) {

Review Comment:
   This sounds like related to #40101, which has been fixed recently



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to