Le 22/09/2026 à 14:30, Jerry D a écrit :
The attached patch is version 2 of 2 of 3.

(Patch 1 of 3 was approved previously and is unchanged.)

This responds to Mikael's comment about using callbacks. Callbacks are eliminated in this version.

Regression tested on x86_64.

OK for mainline?

Regards,

Jerry

---

libgfortran: [PR127347] caf_shmem: do not wait for
  terminated images

When an image terminated normally or with FAIL IMAGE, the other images
kept counting it in their team barriers and blocked forever in the next
SYNC ALL.  An image executing STOP or FAIL IMAGE now drops out of the
barriers of every team it is a member of before it exits.  Lowering the
count wakes the images already waiting once it reaches zero.

An image that terminates without STOP or FAIL IMAGE, e.g. by a signal,
cannot do so.  It now initiates error termination of all images, as for
ERROR STOP, instead of being marked failed.  Only the fork based
supervisor does so.

I'm not sure it is correct.

Failed state is defined as (F2023 5.3.6: Image execution states):
  > An image that has ceased participating in program execution but has
  > not initiated termination is a failed image.

Killed images match this definition. And the definition also kind of implicitly means that the other images continue their execution.

I'm starting to understand what was the reason for the callback in the original patch. The supervisor has no access to the images' teams data, so that the call to leave_teams that is done in the STOP case can't be done by the supervisor in the signal case. Is that correct?

Maybe we can store a copy of caf_current_team->u.image_info and caf_teams_formed->u.image_info in supervisor's image_tracker, so that a call to leave_teams or similar can be done from the supervisor.
A pointer to the parent would also be necessary in shmem_image_info.
Unfortunately it's not very nice to have every image updating rapidly the current team of their image info stored in shared memory, to support something as rare as a signal. I have nothing better to propose so far.

SYNC IMAGES indexed its synchronization table with the team's image
count as the row stride.  That count shrinks once an image has
terminated, although the table is allocated for all images.  Use the
total number of images, and handle an image list and * alike.

Assisted-by: Claude Opus 5

     PR libfortran/127347

libgfortran/ChangeLog:

     * caf/shmem.c (mark_stopped): New function.
     (_gfortran_caf_stop_numeric, _gfortran_caf_stop_str): Call it.
     (_gfortran_caf_fail_image): Call leave_teams.
     * caf/shmem/supervisor.c (supervisor_main_loop): Terminate all
     images when an image terminated without STOP or FAIL IMAGE.
     * caf/shmem/sync.c (sync_table): Use the total number of images as
     the row stride.  Handle an image list and all images alike.
     * caf/shmem/teams_mgmt.c (leave_teams): New function.
     * caf/shmem/teams_mgmt.h (leave_teams): Declare.

gcc/testsuite/ChangeLog:

     * gfortran.dg/coarray/fail_image_sync_1.f90: New test.
     * gfortran.dg/coarray/signal_sync_1.f90: New test.
     * gfortran.dg/coarray/stop_sync_1.f90: New test.
     * gfortran.dg/coarray/sync_images_stopped_1.f90: New test.

---

diff --git a/libgfortran/caf/shmem/teams_mgmt.h 
b/libgfortran/caf/shmem/teams_mgmt.h
index f9da4511128..efeeefc6e24 100644
--- a/libgfortran/caf/shmem/teams_mgmt.h
+++ b/libgfortran/caf/shmem/teams_mgmt.h
@@ -86,6 +86,11 @@ extern caf_shmem_team_t caf_teams_formed;
void update_teams_images (caf_shmem_team_t); +/* Drop this image, which terminated, from the barriers of all teams it is a
+   member of.  */
+
+void leave_teams (void);
+
One nit:
Please mention that the number of finished or failed images has to be updated beforehand for the function to have any effect.

 void check_health (int *, char *, size_t);
#define HEALTH_CHECK(stat, errmsg, errlen) check_health (stat, errmsg, errlen)


OK for the non-signal part (all but supervisor.c and signal_sync_1.f90), with the nit above addressed.

Reply via email to