Describe SECCOMP_FILTER_FLAG_RESTART_BEFORE_RECV in the UAPI header, as SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV is described above it, and say in the documentation that supervisors must tolerate ENOENT from SECCOMP_IOCTL_NOTIF_RECV, which the flag makes routine. Note at the fget_raw() call that O_PATH files are allowed on purpose.
Build tested ARCH=x86_64 defconfig (kernel/seccomp.o and headers) with GCC 16.2.0, and SPHINXDIRS=userspace-api htmldocs. Assisted-by: LLM Signed-off-by: Kees Cook <[email protected]> --- Documentation/userspace-api/seccomp_filter.rst | 8 ++++++++ include/uapi/linux/seccomp.h | 1 + tools/include/uapi/linux/seccomp.h | 1 + kernel/seccomp.c | 1 + 4 files changed, 11 insertions(+) diff --git a/Documentation/userspace-api/seccomp_filter.rst b/Documentation/userspace-api/seccomp_filter.rst index b6875ce54fe2..0a00336d773a 100644 --- a/Documentation/userspace-api/seccomp_filter.rst +++ b/Documentation/userspace-api/seccomp_filter.rst @@ -292,6 +292,14 @@ for calls such as ``close`` where callers do not retry on ``EINTR``. A failed notification receive that resets the notification to its initial state is also eligible for restart. +Each restart withdraws the pending notification, so a supervisor woken by +``poll()``, or blocked in ``ioctl(SECCOMP_IOCTL_NOTIF_RECV)``, may find +nothing left to receive, and the ioctl then fails with ``ENOENT``. This +already happens whenever a notifying process is interrupted before receipt, +but with this flag it becomes routine under repeated signals. Supervisors +should treat ``ENOENT`` from ``SECCOMP_IOCTL_NOTIF_RECV`` as a reason to wait +again, not as an error. + The flag requires ``SECCOMP_FILTER_FLAG_NEW_LISTENER`` and can be used with or without ``SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV``. Using both flags allows handlers to run before receipt and defers non-fatal signals during supervisor diff --git a/include/uapi/linux/seccomp.h b/include/uapi/linux/seccomp.h index 30b76aa48355..3a0ee0551261 100644 --- a/include/uapi/linux/seccomp.h +++ b/include/uapi/linux/seccomp.h @@ -25,6 +25,7 @@ #define SECCOMP_FILTER_FLAG_TSYNC_ESRCH (1UL << 4) /* Received notifications wait in killable state (only respond to fatal signals) */ #define SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (1UL << 5) +/* Restart syscalls interrupted before their notification is received */ #define SECCOMP_FILTER_FLAG_RESTART_BEFORE_RECV (1UL << 6) /* diff --git a/tools/include/uapi/linux/seccomp.h b/tools/include/uapi/linux/seccomp.h index 30b76aa48355..3a0ee0551261 100644 --- a/tools/include/uapi/linux/seccomp.h +++ b/tools/include/uapi/linux/seccomp.h @@ -25,6 +25,7 @@ #define SECCOMP_FILTER_FLAG_TSYNC_ESRCH (1UL << 4) /* Received notifications wait in killable state (only respond to fatal signals) */ #define SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (1UL << 5) +/* Restart syscalls interrupted before their notification is received */ #define SECCOMP_FILTER_FLAG_RESTART_BEFORE_RECV (1UL << 6) /* diff --git a/kernel/seccomp.c b/kernel/seccomp.c index e6e1feed0e7b..3eecba5ccdef 100644 --- a/kernel/seccomp.c +++ b/kernel/seccomp.c @@ -1743,6 +1743,7 @@ static long seccomp_notify_addfd(struct seccomp_filter *filter, if (addfd.newfd && !(addfd.flags & SECCOMP_ADDFD_FLAG_SETFD)) return -EINVAL; + /* Allow O_PATH files, as SCM_RIGHTS and pidfd_getfd() do. */ kaddfd.file = fget_raw(addfd.srcfd); if (!kaddfd.file) return -EBADF; -- 2.55.0

