Describe SECCOMP_FILTER_FLAG_RESTART_BEFORE_RECV in the UAPI header, as
SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV is described above it, and say in
the documentation that supervisors must tolerate ENOENT from
SECCOMP_IOCTL_NOTIF_RECV, which the flag makes routine. Note at the
fget_raw() call that O_PATH files are allowed on purpose.

Build tested ARCH=x86_64 defconfig (kernel/seccomp.o and headers) with
GCC 16.2.0, and SPHINXDIRS=userspace-api htmldocs.

Assisted-by: LLM
Signed-off-by: Kees Cook <[email protected]>
---
 Documentation/userspace-api/seccomp_filter.rst | 8 ++++++++
 include/uapi/linux/seccomp.h                   | 1 +
 tools/include/uapi/linux/seccomp.h             | 1 +
 kernel/seccomp.c                               | 1 +
 4 files changed, 11 insertions(+)

diff --git a/Documentation/userspace-api/seccomp_filter.rst 
b/Documentation/userspace-api/seccomp_filter.rst
index b6875ce54fe2..0a00336d773a 100644
--- a/Documentation/userspace-api/seccomp_filter.rst
+++ b/Documentation/userspace-api/seccomp_filter.rst
@@ -292,6 +292,14 @@ for calls such as ``close`` where callers do not retry on 
``EINTR``.
 A failed notification receive that resets the notification to its initial
 state is also eligible for restart.
 
+Each restart withdraws the pending notification, so a supervisor woken by
+``poll()``, or blocked in ``ioctl(SECCOMP_IOCTL_NOTIF_RECV)``, may find
+nothing left to receive, and the ioctl then fails with ``ENOENT``. This
+already happens whenever a notifying process is interrupted before receipt,
+but with this flag it becomes routine under repeated signals. Supervisors
+should treat ``ENOENT`` from ``SECCOMP_IOCTL_NOTIF_RECV`` as a reason to wait
+again, not as an error.
+
 The flag requires ``SECCOMP_FILTER_FLAG_NEW_LISTENER`` and can be used with
 or without ``SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV``. Using both flags allows
 handlers to run before receipt and defers non-fatal signals during supervisor
diff --git a/include/uapi/linux/seccomp.h b/include/uapi/linux/seccomp.h
index 30b76aa48355..3a0ee0551261 100644
--- a/include/uapi/linux/seccomp.h
+++ b/include/uapi/linux/seccomp.h
@@ -25,6 +25,7 @@
 #define SECCOMP_FILTER_FLAG_TSYNC_ESRCH                (1UL << 4)
 /* Received notifications wait in killable state (only respond to fatal 
signals) */
 #define SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (1UL << 5)
+/* Restart syscalls interrupted before their notification is received */
 #define SECCOMP_FILTER_FLAG_RESTART_BEFORE_RECV        (1UL << 6)
 
 /*
diff --git a/tools/include/uapi/linux/seccomp.h 
b/tools/include/uapi/linux/seccomp.h
index 30b76aa48355..3a0ee0551261 100644
--- a/tools/include/uapi/linux/seccomp.h
+++ b/tools/include/uapi/linux/seccomp.h
@@ -25,6 +25,7 @@
 #define SECCOMP_FILTER_FLAG_TSYNC_ESRCH                (1UL << 4)
 /* Received notifications wait in killable state (only respond to fatal 
signals) */
 #define SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (1UL << 5)
+/* Restart syscalls interrupted before their notification is received */
 #define SECCOMP_FILTER_FLAG_RESTART_BEFORE_RECV        (1UL << 6)
 
 /*
diff --git a/kernel/seccomp.c b/kernel/seccomp.c
index e6e1feed0e7b..3eecba5ccdef 100644
--- a/kernel/seccomp.c
+++ b/kernel/seccomp.c
@@ -1743,6 +1743,7 @@ static long seccomp_notify_addfd(struct seccomp_filter 
*filter,
        if (addfd.newfd && !(addfd.flags & SECCOMP_ADDFD_FLAG_SETFD))
                return -EINVAL;
 
+       /* Allow O_PATH files, as SCM_RIGHTS and pidfd_getfd() do. */
        kaddfd.file = fget_raw(addfd.srcfd);
        if (!kaddfd.file)
                return -EBADF;
-- 
2.55.0


Reply via email to