On 10/6/26 3:22 AM, Mikael Morin wrote:
Le 06/10/2026 à 02:58, Jerry D a écrit :
On 10/5/26 10:04 AM, Jerry D wrote:
See the attached patch. In my campaign to identify caf related bugs I
discovered this one.
Regression tested under heavy loads.
OK for mainline?
Regards,
Jerry
I am working on ways to create LLM assisted patches more human readable. Using
Claude, and I suspect other LLMs, one can set up rules and guidelines to
follow. The attached V2 of the subject patch is essentially identical with an
improved commit message that better explanations of the problem and the fix.
There is a bit of a guide here on how to review all this as well.
How the hang happens: image 1 waits in SYNC IMAGES ([2, 3]) when image
3 stops. Image 2 then executes SYNC IMAGES ([1, 3]), finds image 3
stopped and returns with STAT_STOPPED_IMAGE without synchronizing, so
image 1 is not released. FAIL IMAGE leads to the same hang.
The point to check is the order in sync_table_terminated: under the
table lock it releases the waiting images and takes back their
unmatched counts, and only then stores the new status. Outside the
WIN32 supervisor, which this patch does not change, every change of an
image's status to stopped or failed goes through it.
Reading guide: start with sync.h, the new waiting and aborted fields
and what sync_table and sync_table_terminated promise. Then those two
functions in sync.c. In shmem.c, _gfortran_caf_sync_images drops its
early return for a terminated image and sets STAT= only when sync_table
reports a terminated image. The other shmem.c and supervisor.c hunks
route the status changes through sync_table_terminated, which replaces
the supervisor's wake-up of all waiting images. set_table_parts and
the sync_init changes only lay out the enlarged table.
Regression tested on x86_64.
OK for mainline?
Regards,
Jerry
PS Feedback on the format of this email is greatly appreciated.
---
libgfortran: [PR127684] caf_shmem: SYNC IMAGES hang on
terminated image
SYNC IMAGES hung when an image of the image set stopped or failed while
another image of the set was waiting in the statement.
An image that found a stopped or failed image in its set returned at once
with STAT=, without synchronizing, so an image already waiting for it was
never released. The supervisor also woke the waiting images without
holding the table lock, so a wakeup could be lost.
Each image now records the set it waits for. Images that stop or fail,
and the POSIX supervisor when it reaps one, set the status through the
new sync_table_terminated, which holds the table lock. For a stopped
image it releases the images waiting for it and takes back the counts
they raised that no partner has matched, before it stores the status,
so an image that sees the stop cannot pair with the aborted statement.
A failed image is skipped and the other images of the set synchronize
(F2023 11.7.11).
Assisted-by: Claude Opus 5.5
PR libfortran/127684
libgfortran/ChangeLog:
* caf/shmem.c (_gfortran_caf_sync_images): Do not return early for
a stopped or failed image. Check the images only when one of them
terminated.
(mark_stopped): Use sync_table_terminated.
(_gfortran_caf_fail_image): Likewise.
* caf/shmem/supervisor.c (supervisor_main_loop): Likewise. Do not
signal all images waiting in sync_table.
* caf/shmem/sync.c (set_table_parts): New function.
(sync_init): Use it.
(sync_init_supervisor): Allocate the image sets of the waiting
images and the aborted flags.
(stopped_image_in): New function.
(sync_table): Return whether an image of the set terminated. Only
synchronize memory when an image of the set stopped. Record the
image set and end the wait when aborted.
(sync_table_terminated): New function.
* caf/shmem/sync.h (sync_t): Add waiting and aborted.
(sync_table): Return bool.
(sync_table_terminated): Declare.
gcc/testsuite/ChangeLog:
* gfortran.dg/coarray/sync_images_failed_1.f90: New test.
* gfortran.dg/coarray/sync_images_stopped_2.f90: New test.
---
+void
+sync_table_terminated (sync_t *si, int image, bool stopped)
As far as I can see, the function doesn't impact only sync_table as it also
updates the supervisor tables, so please remove sync_table from the name
(notify_image_terminated, image_terminated, acknowledge_image_terminated or
similar).
+{
+ volatile int *table = si->table;
+ const size_t img_c = local->total_num_images;
+
+ lock_table (si);
+ for (size_t j = 0; j < img_c; ++j)
+ {
+ if (!si->waiting[image + img_c * j])
+ continue;
+ if (stopped)
+ {
+ /* Take back the counts of J that its partners have not matched.
+ IMAGE is marked stopped only below, so no image that sees the
+ stop can pair with J's aborted SYNC IMAGES. */
+ for (size_t k = 0; k < img_c; ++k)
+ if (si->waiting[k + img_c * j]
+ && table[k + img_c * j] > table[j + img_c * k])
+ --table[k + img_c * j];
It took me some time to understand why this loop was necessary. I finally
convinced myself that it makes sense, but I think we can do without it if...
+ si->aborted[j] = 1;
+ }
+ caf_shmem_cond_signal (&si->triggers[j]);
+ }
+ /* The images woken above need the table lock to return from their wait,>
+ so they see this count and the new status. */
+ atomic_fetch_add (stopped ? &this_image.supervisor->finished_images
+ : &this_image.supervisor->failed_images, 1);
+ this_image.supervisor->images[image].status
+ = stopped ? IMAGE_SUCCESS : IMAGE_FAILED;
unlock_table (si);
}
@@ -106,6 +131,16 @@ sync_table (sync_t *si, int *images, int size)
}
lock_table (si);
+ /* With a stopped image in the set this only has the effect of SYNC
+ MEMORY; failed images are skipped (F2023 11.7.11). */
+ if (stopped_image_in (images, size))
+ {
+ unlock_table (si);
+ return true;
+ }
... this is moved further down...
+ for (i = 0; i < size; ++i)
+ si->waiting[images[i] + img_c * this_image.image_num] = 1;
+ si->aborted[this_image.image_num] = 0;
for (i = 0; i < size; ++i)
{
if (this_image.supervisor->images[images[i]].status != IMAGE_OK)
... after this loop. Instead of decreasing the count of already arrived images,
we would not skip the increase of the count as images arrive to the barrier.
@@ -115,6 +150,9 @@ sync_table (sync_t *si, int *images, int size)
}
for (;;)
{
+ /* An image of the set stopped, see sync_table_terminated. */
+ if (si->aborted[this_image.image_num])
+ break;
The information present in the aborted array seems to be redundant with the
status of images shared with the supervisor.
for (i = 0; i < size; ++i)
if (this_image.supervisor->images[images[i]].status == IMAGE_OK
As the status is already queried in the main loop, we can just as well handle
status different from IMAGE_OK here, so that the aborted array is no longer
necessary.
&& table[images[i] + img_c * this_image.image_num]
Next there is the waiting array that I would like to remove as well.
As far as I can see, with the count decrease loop removed and the aborted array
removed, the waiting array remains useful to selectively wake the needed images
in sync_table_terminated. But it requires updating it in every sync image, just
to support a premature image stop. Can we just wake every image in the team on
image stop and remove the waiting array and the associated book keeping in the
main synchronisation code?
The explaining text is an improvement I think, but I had already a clear picture
of the patch in mind when I started reading it. Let's see how it helps with a
completely unknown patch.
Thanks for your work.
Hi Mikael,
Thanks for the review. The attached v3 takes all of your suggestions,
and the patch is much smaller for it.
Changes since v2:
- sync_table_terminated is renamed notify_image_terminated.
- sync_table counts the statement for every active image of the set
before it checks for a stopped image, so the counts of the images
that remain active always agree. The loop that took back unmatched
counts is gone.
- The wait loop ends when the status of an image of the set is stopped,
so the aborted array is gone.
- notify_image_terminated wakes every image, so the waiting array is
gone as well. The table is back to N x N and set_table_parts is
gone.
Counting the statement also changes one behaviour, for the better I
think. F2023 11.7.4 pairs SYNC IMAGES statements by how many times
each image has executed one with the other image in its set, and a
statement that finds a stopped image has still been executed. v2 did
not count it, so when image 1 is in SYNC IMAGES (2) and image 2 runs
SYNC IMAGES ([1, 3]) after image 3 stopped, image 1 kept waiting for a
later statement of image 2. The new test sync_images_stopped_3.f90
covers this case; it hangs with v2.
Reading guide: start with sync_table in sync.c, the count loop and the
stopped-image check at the top of the wait loop, then
notify_image_terminated below it. In shmem.c,
_gfortran_caf_sync_images drops its early return for a terminated
image and sets STAT= only when sync_table reports one. The other
shmem.c and supervisor.c hunks route the status changes through
notify_image_terminated, which replaces the supervisor's wake-up
without the table lock.
Note: The WIN32 path is not changed. I do not have a Wondows based system to
test with at the moment. It will have to be a later smaller patch.
Tested with the three new tests, and with a matrix of 66 cases (STOP,
FAIL IMAGE and a killing signal against SYNC IMAGES with a list and
with *, in the initial team and in a child team), 500 runs quiet and
500 under full CPU load: no hangs, and the right STAT= in every run.
Regression tested on x86_64.
OK for mainline?
Regards,
Jerry
From b69d7d7842e64dafa39bad8d7aab6861f9824f86 Mon Sep 17 00:00:00 2001
From: Jerry DeLisle <[email protected]>
Date: Wed, 30 Sep 2026 19:26:47 -0700
Subject: [PATCH v3] libgfortran: [PR127684] caf_shmem: SYNC IMAGES hang on
terminated image
SYNC IMAGES hung when an image of the image set stopped or failed while
another image of the set was waiting in the statement.
An image that found a stopped or failed image in its set returned at once
with STAT=, without counting the statement for the other images of the
set, so an image already waiting for it was never released. The
supervisor, the parent process that starts the images and waits for them
to exit, also woke the waiting images without holding the table lock, so
a wakeup could be lost.
Each image now counts the statement for every active image of its set,
also when one of them has stopped, so that corresponding statements keep
pairing (F2023 11.7.4). With a stopped image the statement then only has
the effect of SYNC MEMORY, and failed images are skipped (F2023 11.7.11).
Every change of an image's status to stopped or failed now goes through
the new notify_image_terminated, which holds the table lock and wakes
every image waiting in SYNC IMAGES. The image sets its own status in
STOP or FAIL IMAGE; the supervisor sets it when an image exited or was
killed by a signal without doing so. The Windows version of the
supervisor is not changed.
Assisted-by: Claude Opus 5.5
PR libfortran/127684
libgfortran/ChangeLog:
* caf/shmem.c (_gfortran_caf_sync_images): Do not return early for
a stopped or failed image. Check the images only when one of them
terminated.
(mark_stopped): Use notify_image_terminated.
(_gfortran_caf_fail_image): Likewise.
* caf/shmem/supervisor.c (supervisor_main_loop): Likewise. Do not
signal the images waiting in sync_table.
* caf/shmem/sync.c (stopped_image_in): New function.
(sync_table): Return whether an image of the set terminated. Stop
waiting when an image of the set stopped.
(notify_image_terminated): New function.
* caf/shmem/sync.h (sync_table): Return bool.
(notify_image_terminated): Declare.
gcc/testsuite/ChangeLog:
* gfortran.dg/coarray/sync_images_failed_1.f90: New test.
* gfortran.dg/coarray/sync_images_stopped_2.f90: New test.
* gfortran.dg/coarray/sync_images_stopped_3.f90: New test.
---
.../coarray/sync_images_failed_1.f90 | 55 ++++++++++++++++++
.../coarray/sync_images_stopped_2.f90 | 58 +++++++++++++++++++
.../coarray/sync_images_stopped_3.f90 | 40 +++++++++++++
libgfortran/caf/shmem.c | 34 +++--------
libgfortran/caf/shmem/supervisor.c | 11 +---
libgfortran/caf/shmem/sync.c | 40 ++++++++++++-
libgfortran/caf/shmem/sync.h | 11 +++-
7 files changed, 212 insertions(+), 37 deletions(-)
create mode 100644 gcc/testsuite/gfortran.dg/coarray/sync_images_failed_1.f90
create mode 100644 gcc/testsuite/gfortran.dg/coarray/sync_images_stopped_2.f90
create mode 100644 gcc/testsuite/gfortran.dg/coarray/sync_images_stopped_3.f90
diff --git a/gcc/testsuite/gfortran.dg/coarray/sync_images_failed_1.f90 b/gcc/testsuite/gfortran.dg/coarray/sync_images_failed_1.f90
new file mode 100644
index 00000000000..c076ee4d88e
--- /dev/null
+++ b/gcc/testsuite/gfortran.dg/coarray/sync_images_failed_1.f90
@@ -0,0 +1,55 @@
+! { dg-do run }
+! PR 127684
+!
+! SYNC IMAGES (STAT=) with a failed image in the image set synchronizes the
+! active images of the set (F2023 11.7.11), both for an image that already
+! waits in the statement when the image fails and for one that arrives later.
+!
+! Image 1 tells image 3 that it is about to wait for images 2 and 3; image 3
+! spins to give it time to start waiting, then fails. Image 2 executes its
+! SYNC IMAGES only after it sees image 3 failed, and image 1 reads X from
+! image 2 to check that the two synchronized. Without the fix the test hangs.
+
+program sync_images_failed_1
+ use iso_fortran_env, only : atomic_int_kind, stat_failed_image
+ implicit none
+ integer(atomic_int_kind) :: arrived[*], a
+ integer :: st, x[*]
+
+ if (num_images () < 3) stop
+ arrived = 0
+ x = 0
+ sync all
+
+ select case (this_image ())
+ case (1)
+ call atomic_define (arrived[3], 1)
+ st = 0
+ sync images ([2, 3], stat=st)
+ if (st /= stat_failed_image) stop 1
+ if (x[2] /= 42) stop 2
+ case (2)
+ do while (image_status (3) /= stat_failed_image)
+ end do
+ x = 42
+ st = 0
+ sync images ([1, 3], stat=st)
+ if (st /= stat_failed_image) stop 3
+ case (3)
+ a = 0
+ do while (a == 0)
+ call atomic_ref (a, arrived)
+ end do
+ call spin ()
+ fail image
+ end select
+contains
+ subroutine spin ()
+ integer :: c
+ integer(kind=8) :: v
+ v = 2
+ do c = 1, 20000000
+ v = mod (v * 2, 199679_8)
+ end do
+ end subroutine
+end program sync_images_failed_1
diff --git a/gcc/testsuite/gfortran.dg/coarray/sync_images_stopped_2.f90 b/gcc/testsuite/gfortran.dg/coarray/sync_images_stopped_2.f90
new file mode 100644
index 00000000000..5d5ac61c83f
--- /dev/null
+++ b/gcc/testsuite/gfortran.dg/coarray/sync_images_stopped_2.f90
@@ -0,0 +1,58 @@
+! { dg-do run }
+! PR 127684
+!
+! SYNC IMAGES (STAT=) with a stopped image in the image set has the effect of
+! SYNC MEMORY (F2023 11.7.11), both for an image that already waits in the
+! statement when the image stops and for one that arrives later. The next
+! SYNC IMAGES of these two images has to pair with each other.
+!
+! Image 1 tells image 3 that it is about to wait for images 2 and 3; image 3
+! spins to give it time to start waiting, then stops. Image 2 executes its
+! first SYNC IMAGES only after it sees image 3 stopped, and image 1 checks
+! the pairing of the next one through X. Without the fix the test hangs.
+
+program sync_images_stopped_2
+ use iso_fortran_env, only : atomic_int_kind, stat_stopped_image
+ implicit none
+ integer(atomic_int_kind) :: arrived[*], a
+ integer :: st, x[*]
+
+ if (num_images () < 3) stop
+ arrived = 0
+ x = 0
+ sync all
+
+ select case (this_image ())
+ case (1)
+ call atomic_define (arrived[3], 1)
+ st = 0
+ sync images ([2, 3], stat=st)
+ if (st /= stat_stopped_image) stop 1
+ sync images (2)
+ if (x[2] /= 42) stop 2
+ case (2)
+ do while (image_status (3) /= stat_stopped_image)
+ end do
+ st = 0
+ sync images ([1, 3], stat=st)
+ if (st /= stat_stopped_image) stop 3
+ x = 42
+ sync images (1)
+ case (3)
+ a = 0
+ do while (a == 0)
+ call atomic_ref (a, arrived)
+ end do
+ call spin ()
+ stop
+ end select
+contains
+ subroutine spin ()
+ integer :: c
+ integer(kind=8) :: v
+ v = 2
+ do c = 1, 20000000
+ v = mod (v * 2, 199679_8)
+ end do
+ end subroutine
+end program sync_images_stopped_2
diff --git a/gcc/testsuite/gfortran.dg/coarray/sync_images_stopped_3.f90 b/gcc/testsuite/gfortran.dg/coarray/sync_images_stopped_3.f90
new file mode 100644
index 00000000000..f479a7a0196
--- /dev/null
+++ b/gcc/testsuite/gfortran.dg/coarray/sync_images_stopped_3.f90
@@ -0,0 +1,40 @@
+! { dg-do run }
+! PR 127684
+!
+! A SYNC IMAGES that finds a stopped image in its set still counts for the
+! correspondence with the other images of the set (F2023 11.7.4), so it
+! releases an image whose set does not contain the stopped image.
+!
+! Image 2 executes its first SYNC IMAGES only after it sees image 3 stopped;
+! image 1 may wait for it already or arrive later. Image 1 checks the
+! pairing of the next one through X. Without the fix image 1 pairs its first
+! SYNC IMAGES with the second one of image 2 and the test hangs.
+
+program sync_images_stopped_3
+ use iso_fortran_env, only : stat_stopped_image
+ implicit none
+ integer :: st, x[*]
+
+ if (num_images () < 3) stop
+ x = 0
+ sync all
+
+ select case (this_image ())
+ case (1)
+ st = -1
+ sync images (2, stat=st)
+ if (st /= 0) stop 1
+ sync images (2)
+ if (x[2] /= 42) stop 2
+ case (2)
+ do while (image_status (3) /= stat_stopped_image)
+ end do
+ st = 0
+ sync images ([1, 3], stat=st)
+ if (st /= stat_stopped_image) stop 3
+ x = 42
+ sync images (1)
+ case (3)
+ stop
+ end select
+end program sync_images_stopped_3
diff --git a/libgfortran/caf/shmem.c b/libgfortran/caf/shmem.c
index 08ae3113da5..7fb23f05e6c 100644
--- a/libgfortran/caf/shmem.c
+++ b/libgfortran/caf/shmem.c
@@ -528,23 +528,6 @@ _gfortran_caf_sync_images (int count, int images[], int *stat, char *errmsg,
if (images[c] > 0 && images[c] <= max_id)
{
mapped_images[c] = map[images[c] - 1];
- switch (this_image.supervisor->images[mapped_images[c]].status)
- {
- case IMAGE_SUCCESS:
- caf_internal_error ("SYNC IMAGES: Image %d is stopped", stat,
- errmsg, errmsg_len, images[c]);
- /* We can come here only, when stat is non-NULL. */
- *stat = CAF_STAT_STOPPED_IMAGE;
- return;
- case IMAGE_FAILED:
- caf_internal_error ("SYNC IMAGES: Image %d has failed", stat,
- errmsg, errmsg_len, images[c]);
- /* We can come here only, when stat is non-NULL. */
- *stat = CAF_STAT_FAILED_IMAGE;
- return;
- default:
- break;
- }
for (int i = 0; i < c; ++i)
if (mapped_images[c] == mapped_images[i])
{
@@ -570,8 +553,12 @@ _gfortran_caf_sync_images (int count, int images[], int *stat, char *errmsg,
HEALTH_CHECK (stat, errmsg, errmsg_len);
__asm__ __volatile__ ("" ::: "memory");
- sync_table (&local->si, mapped_images, count);
- if (count > 0)
+ if (!sync_table (&local->si, mapped_images, count))
+ {
+ if (stat)
+ *stat = 0;
+ }
+ else if (count > 0)
check_health (mapped_images, count, stat, errmsg, errmsg_len);
else
HEALTH_CHECK (stat, errmsg, errmsg_len);
@@ -589,11 +576,7 @@ mark_stopped (void)
return;
if (this_image.supervisor->images[this_image.image_num].status == IMAGE_OK)
- {
- this_image.supervisor->images[this_image.image_num].status
- = IMAGE_SUCCESS;
- atomic_fetch_add (&this_image.supervisor->finished_images, 1);
- }
+ notify_image_terminated (&local->si, this_image.image_num, true);
leave_teams (true);
}
@@ -657,8 +640,7 @@ void
_gfortran_caf_fail_image (void)
{
fputs ("IMAGE FAILED!\n", stderr);
- this_image.supervisor->images[this_image.image_num].status = IMAGE_FAILED;
- atomic_fetch_add (&this_image.supervisor->failed_images, 1);
+ notify_image_terminated (&local->si, this_image.image_num, false);
leave_teams (false);
exit (0);
}
diff --git a/libgfortran/caf/shmem/supervisor.c b/libgfortran/caf/shmem/supervisor.c
index 246f98aab63..41e2c160063 100644
--- a/libgfortran/caf/shmem/supervisor.c
+++ b/libgfortran/caf/shmem/supervisor.c
@@ -432,8 +432,7 @@ supervisor_main_loop (int *argc __attribute__ ((unused)),
image already. */
if (m->images[j].status == IMAGE_OK)
{
- m->images[j].status = IMAGE_SUCCESS;
- atomic_fetch_add (&m->finished_images, 1);
+ notify_image_terminated (&local->si, j, true);
}
}
else if (!WIFEXITED (chstatus) || WEXITSTATUS (chstatus))
@@ -484,8 +483,7 @@ supervisor_main_loop (int *argc __attribute__ ((unused)),
already, e.g. by a STOP with a non-zero stop code. */
if (m->images[j].status == IMAGE_OK)
{
- m->images[j].status = IMAGE_FAILED;
- atomic_fetch_add (&m->failed_images, 1);
+ notify_image_terminated (&local->si, j, false);
/* The image did not leave the barriers of its teams, so do
it for it. */
update_registered_teams ();
@@ -496,11 +494,6 @@ supervisor_main_loop (int *argc __attribute__ ((unused)),
*exit_code = 1;
}
}
- /* Trigger waiting sync images aka sync_table. */
- for (j = 0; j < local->total_num_images; j++)
- caf_shmem_cond_signal (&SHMPTR_AS (caf_shmem_condvar *,
- m->sync_shared.sync_images_cond_vars,
- &local->sm)[j]);
counter_barrier_add (&m->num_active_images, -1);
#elif defined(WIN32)
DWORD res = WaitForMultipleObjects (count_waiting, waiting_handles, FALSE,
diff --git a/libgfortran/caf/shmem/sync.c b/libgfortran/caf/shmem/sync.c
index 9deb31f0871..e3041a3c128 100644
--- a/libgfortran/caf/shmem/sync.c
+++ b/libgfortran/caf/shmem/sync.c
@@ -81,7 +81,18 @@ sync_init_supervisor (sync_t *si, alloc *ai)
memset (si->table, 0, table_size_in_bytes);
}
-void
+/* Whether one of the SIZE IMAGES has stopped. */
+
+static bool
+stopped_image_in (const int *images, int size)
+{
+ for (int i = 0; i < size; ++i)
+ if (this_image.supervisor->images[images[i]].status == IMAGE_SUCCESS)
+ return true;
+ return false;
+}
+
+bool
sync_table (sync_t *si, int *images, int size)
{
/* The variable `table` is an N x N matrix, where N is the number of all
@@ -97,6 +108,7 @@ sync_table (sync_t *si, int *images, int size)
/* The table is allocated for all images, so the row stride is the total
number of images and not the (shrinking) number of images in a team. */
const size_t img_c = local->total_num_images;
+ bool terminated;
int i;
if (size <= 0)
@@ -106,6 +118,8 @@ sync_table (sync_t *si, int *images, int size)
}
lock_table (si);
+ /* Count the statement even with a stopped image in the set, so that the
+ counts of the active images keep pairing corresponding statements. */
for (i = 0; i < size; ++i)
{
if (this_image.supervisor->images[images[i]].status != IMAGE_OK)
@@ -115,6 +129,10 @@ sync_table (sync_t *si, int *images, int size)
}
for (;;)
{
+ /* With a stopped image in the set this only has the effect of SYNC
+ MEMORY; failed images are skipped (F2023 11.7.11). */
+ if (stopped_image_in (images, size))
+ break;
for (i = 0; i < size; ++i)
if (this_image.supervisor->images[images[i]].status == IMAGE_OK
&& table[images[i] + img_c * this_image.image_num]
@@ -125,6 +143,26 @@ sync_table (sync_t *si, int *images, int size)
caf_shmem_cond_wait (&si->triggers[this_image.image_num],
&si->cis->sync_images_table_lock);
}
+ terminated = false;
+ for (i = 0; i < size; ++i)
+ if (this_image.supervisor->images[images[i]].status != IMAGE_OK)
+ terminated = true;
+ unlock_table (si);
+ return terminated;
+}
+
+void
+notify_image_terminated (sync_t *si, int image, bool stopped)
+{
+ /* Under the table lock, so that an image about to wait in sync_table
+ either sees the new status or is woken. */
+ lock_table (si);
+ atomic_fetch_add (stopped ? &this_image.supervisor->finished_images
+ : &this_image.supervisor->failed_images, 1);
+ this_image.supervisor->images[image].status
+ = stopped ? IMAGE_SUCCESS : IMAGE_FAILED;
+ for (int j = 0; j < local->total_num_images; ++j)
+ caf_shmem_cond_signal (&si->triggers[j]);
unlock_table (si);
}
diff --git a/libgfortran/caf/shmem/sync.h b/libgfortran/caf/shmem/sync.h
index d7b96992840..b359c17cbe3 100644
--- a/libgfortran/caf/shmem/sync.h
+++ b/libgfortran/caf/shmem/sync.h
@@ -69,7 +69,16 @@ void sync_team (caf_shmem_team_t team);
bool sync_team_unless_stopped (caf_shmem_team_t team, int *terminated);
-void sync_table (sync_t *, int *, int);
+/* Synchronize with the SIZE images in the array, or with the images of the
+ current team when SIZE is zero. Returns true when an image of the set
+ terminated before the images synchronized. */
+
+bool sync_table (sync_t *, int *, int);
+
+/* Set the status of IMAGE, which terminated, to stopped or failed, count it
+ and wake the images waiting in sync_table. */
+
+void notify_image_terminated (sync_t *, int image, bool stopped);
void lock_alloc_lock (sync_t *);
--
2.55.0