Le 30/09/2026 à 00:17, Jerry D a écrit :
On 9/29/26 1:14 AM, Mikael Morin wrote:
Le 28/09/2026 à 22:30, Jerry D a écrit :
I have been out on travel this weekend. I am curious why the email is
marked SPAM? I will look at your comments tomorro.
Yes, sorry about it, I forgot to remove the SPAM marker. There is
obviously a spam filter before I receive the message; it selects its
messages more or less randomly, and it marks them with a tag in the
subject rather than with a dedicated attribute. Yes, it's more a
nuisance than an help, but it's out of my control.
Attached is v2 of the patch. I had claude run fuzz tests on this which
generated about 900 variations. Doing this identified another small
adjustment needed to avoid some timing issues under heavy loads in the
counter_barrier.c code.
Regression tested on x86_64.
Ok for mainline?
Regards,
Jerry
---
libgfortran: [PR127347] caf_shmem: drop images killed by a
signal from teams
An image that terminates without STOP or FAIL IMAGE, e.g. when it is
killed by a signal, cannot drop itself from the barriers of its teams,
so the other images blocked forever in their next image control
statement. Only the supervisor notices such an image, and it has no
access to the teams the image was a member of.
Every team now registers itself with the supervisor when it is formed,
which happens once per team, so that the supervisor can walk the teams
and do for a killed image what the image itself does for a STOP.
Entries are pushed onto the list without a lock, so that an image
killed while registering a team cannot block the supervisor.
An image killed this way stays a failed image (F2023 5.3.6): the other
images of its teams carry on and report it through STAT=.
Reporting it needs SYNC ALL and SYNC TEAM to check the images of the
team after synchronizing. They only did so when they had not
synchronized, which is the case for a stopped image; with a failed
image the active images are synchronized (F2023 11.7.11) and the only
other check ran before the barrier, so an image failing while the
others already waited in the statement left STAT= zero.
Assisted-by: Claude Opus 5.5
PR libfortran/127347
libgfortran/ChangeLog:
* caf/shmem.c (_gfortran_caf_init): Synchronize without the checks
of the SYNC ALL statement.
(_gfortran_caf_sync_all, _gfortran_caf_sync_team): Report the images
of the team that terminated before the images synchronized.
(_gfortran_caf_form_team): Register the team.
* caf/shmem/collective_subroutine.c (collsub_sync): Adjust.
* caf/shmem/counter_barrier.c (counter_barrier_init): Initialize the
new fields.
(counter_barrier_add_locked): Count the images removed from the
barrier.
(next_round): Take that count as the round ends.
(wait_round, counter_barrier_wait_abortable): Report it for the round
the image took part in.
* caf/shmem/counter_barrier.h (counter_barrier): Add terminated,
round_terminated and round_terminated_round.
(counter_barrier_wait_abortable): Adjust.
* caf/shmem/supervisor.c (ensure_shmem_initialization): Initialize
the list of teams. Register the initial team.
(supervisor_main_loop): Drop an image that terminated without
leaving its teams from them.
* caf/shmem/supervisor.h (supervisor): Add teams.
* caf/shmem/sync.c (sync_all, sync_team_unless_stopped): Report the
images that terminated.
* caf/shmem/sync.h (sync_all, sync_team_unless_stopped): Adjust.
* caf/shmem/teams_mgmt.c (team_terminated_images)
(update_teams_images_locked, leave_team): Take the team's image
info.
(register_team, update_registered_teams): New functions.
* caf/shmem/teams_mgmt.h (shmem_image_info): Add next_team.
(register_team, update_registered_teams): Declare.
gcc/testsuite/ChangeLog:
* gfortran.dg/coarray/signal_sync_1.f90: New test.
* gfortran.dg/coarray/signal_team_1.f90: New test.
* gfortran.dg/coarray/sync_failed_stat_1.f90: New test.
---
Hello,
there is just one thing that is not completely clear to me.
diff --git a/libgfortran/caf/shmem/counter_barrier.c
b/libgfortran/caf/shmem/counter_barrier.c
index 8598d79d69d..43e80289d23 100644
--- a/libgfortran/caf/shmem/counter_barrier.c
+++ b/libgfortran/caf/shmem/counter_barrier.c
@@ -66,6 +69,10 @@ counter_barrier_init (counter_barrier *b, int val)
static void
next_round (counter_barrier *b, bool abort)
{
+ /* Take the images removed so far as the ones that were involved in the
+ round; an image removed later did not take part in it. */
+ b->round_terminated = b->terminated;
+ b->round_terminated_round = b->curr_wait_group;
if (abort)
b->aborted_round = b->curr_wait_group;
++b->curr_wait_group;
@@ -78,9 +85,10 @@ next_round (counter_barrier *b, bool abort)
false, when the round was aborted. */
static bool
-wait_round (counter_barrier *b, bool abortable)
+wait_round (counter_barrier *b, bool abortable, int *terminated)
{
const uint64_t round = b->curr_wait_group;
+ const int at_arrival = b->terminated;
if (abortable)
++b->abortable_arrivals;
@@ -93,6 +101,12 @@ wait_round (counter_barrier *b, bool abortable)
if (b->curr_wait_group == round)
next_round (b, false);
+ if (terminated)
+ /* A later round may have taken its own snapshot in the meantime; the
+ images removed before this one arrived are known to be involved. */
+ *terminated = b->round_terminated_round == round ? b->round_terminated
+ : at_arrival;
+
It seems that returning at_arrival, which is the value saved before
waiting, would return an obsolete value. In which case is
b->round_terminated_round different from round? Isn't next_round run
just once before each image executes this?
return b->aborted_round != round;
}