> diff --git a/tools/testing/selftests/bpf/progs/refcounted_kptr.c
> b/tools/testing/selftests/bpf/progs/refcounted_kptr.c
> index fd35093285c0d..383c5b1b7111f 100644
> --- a/tools/testing/selftests/bpf/progs/refcounted_kptr.c
> +++ b/tools/testing/selftests/bpf/progs/refcounted_kptr.c
> @@ -893,6 +893,106 @@ long refcount_acquire_rcu_map_kptr_null_checked(void
> *ctx)
> return 0;
> }
>
> +SEC("?syscall")
> +__success
> +long map_kptr_read_after_rcu_unlock(void *ctx)
A reviewer noted in v3 that the tests use TC programs with implicit RCU
protection, causing pointers not to become untrusted as intended, and
specifically suggested: "The right fix might be to use fentry.s for these
tests to trigger untrusted marking of pointers."
v4 uses SEC("?syscall") instead. While syscall programs are sleepable and
may trigger untrusted marking, does this section annotation achieve the
same behavior Kumar recommended fentry.s for? The deviation from the
specific suggestion wasn't discussed in the v4 thread.
> +{
> + struct map_value_refcount_only *mapval;
> + struct node_refcount_only *n;
> + int idx = 0;
> +
> + mapval = bpf_map_lookup_elem(&stashed_refcount_only, &idx);
> + if (!mapval)
> + return 0;
> +
> + bpf_rcu_read_lock();
> + n = mapval->node;
> + if (!n) {
> + bpf_rcu_read_unlock();
> + return 0;
> + }
> + bpf_rcu_read_unlock();
> +
> + return n->key;
> +}
> +
> +SEC("?syscall")
> +__failure __msg("is neither owning or non-owning ref")
> +long refcount_acquire_graph_after_rcu_unlock(void *ctx)
> +{
> + struct map_value *mapval;
> + struct node_data *n, *m;
> + int idx = 0;
> +
> + mapval = bpf_map_lookup_elem(&stashed_nodes, &idx);
> + if (!mapval)
> + return 0;
> +
> + bpf_rcu_read_lock();
> + n = mapval->node;
> + if (!n) {
> + bpf_rcu_read_unlock();
> + return 0;
> + }
> + bpf_rcu_read_unlock();
> +
> + m = bpf_refcount_acquire(n);
> + if (m)
> + bpf_obj_drop(m);
> +
> + return 0;
> +}
> +
> +SEC("?syscall")
> +__failure __msg("only read is supported")
> +long graph_map_kptr_write_after_rcu_unlock(void *ctx)
> +{
> + struct map_value *mapval;
> + struct node_data *n;
> + int idx = 0;
> +
> + mapval = bpf_map_lookup_elem(&stashed_nodes, &idx);
> + if (!mapval)
> + return 1;
> +
> + bpf_rcu_read_lock();
> + n = mapval->node;
> + if (!n) {
> + bpf_rcu_read_unlock();
> + return 2;
> + }
> + bpf_rcu_read_unlock();
> +
> + n->key = 1;
> + return 0;
> +}
Does this test actually verify the fix? The "only read is supported"
rejection comes from the PTR_UNTRUSTED check in check_ptr_to_btf_access():
if (atype != BPF_READ && (type_flag(reg->type) & PTR_UNTRUSTED)) {
verbose(env, "only read is supported\n");
return -EACCES;
}
This check is unchanged in the pre-series tree and rejects writes through
any PTR_UNTRUSTED register for both object layouts. A kernel missing the
series fix would still fail this test with the same message, so the test
wouldn't catch a revert.
> +
> +SEC("?syscall")
> +__success
> +long graph_map_kptr_read_after_spin_unlock(void *ctx)
> +{
> + struct map_value *mapval;
> + struct node_data *n;
> + int idx = 0;
> +
> + mapval = bpf_map_lookup_elem(&stashed_nodes, &idx);
> + if (!mapval)
> + return 0;
> +
> + bpf_rcu_read_lock();
> + n = mapval->node;
> + if (!n) {
> + bpf_rcu_read_unlock();
> + return 0;
> + }
> + bpf_rcu_read_unlock();
> +
> + bpf_spin_lock(&lock);
> + bpf_spin_unlock(&lock);
> +
> + return n->key;
> +}
Is __success correct for this read? After bpf_rcu_read_unlock() in a
sleepable program, the pointer gains PTR_UNTRUSTED but keeps NON_OWN_REF
because the fix excludes PTR_UNTRUSTED from type_is_ptr_alloc_obj().
At the load, the type is PTR_TO_BTF_ID | MEM_ALLOC | PTR_UNTRUSTED |
NON_OWN_REF. The bpf_convert_ctx_accesses() switch in
kernel/bpf/fixups.c lists these BPF_PROBE_MEM cases:
case PTR_TO_BTF_ID | PTR_UNTRUSTED:
case PTR_TO_BTF_ID | MEM_ALLOC | PTR_UNTRUSTED:
but not the NON_OWN_REF variant, so the switch falls through and the load
stays a plain BPF_MEM access with no exception-table entry.
The object is freed via bpf_mem_free_rcu(), so in a sleepable syscall
program that has left the RCU critical section, a grace period can elapse
and the page can be returned to the allocator before the load executes,
causing an unprotected fault.
Pre-series, invalidate_non_owning_refs() at bpf_spin_unlock() killed this
register because type_is_non_owning_ref() was true. Now that
type_is_ptr_alloc_obj() excludes PTR_UNTRUSTED, type_is_non_owning_ref()
returns false and the stale register survives.
The sibling test map_kptr_read_after_rcu_unlock (node_refcount_only, no
graph node, hence no NON_OWN_REF) does match a BPF_PROBE_MEM case and gets
the probe rewrite. Should either fixups.c gain a NON_OWN_REF case and this
test document the probe path, or should NON_OWN_REF be stripped in
invalidate_rcu_protected_refs()?
> +
> static long __stash_map_empty_xchg(struct node_data *n, int idx)
> {
> struct map_value *mapval = bpf_map_lookup_elem(&stashed_nodes, &idx);
> diff --git a/tools/testing/selftests/bpf/progs/refcounted_kptr_fail.c
> b/tools/testing/selftests/bpf/progs/refcounted_kptr_fail.c
> index acd3e81a39168..0cc4cbd0c81b1 100644
> --- a/tools/testing/selftests/bpf/progs/refcounted_kptr_fail.c
> +++ b/tools/testing/selftests/bpf/progs/refcounted_kptr_fail.c
> @@ -127,6 +127,33 @@ long refcount_acquire_rcu_map_kptr_unchecked_drop(void
> *ctx)
> return 0;
> }
>
> +SEC("?syscall")
> +__failure __msg("is neither owning or non-owning ref")
> +long refcount_acquire_after_rcu_unlock(void *ctx)
> +{
> + struct map_value_refcount_only *mapval;
> + struct node_refcount_only *n, *m;
> + int idx = 0;
> +
> + mapval = bpf_map_lookup_elem(&stashed_refcount_only, &idx);
> + if (!mapval)
> + return 1;
> +
> + bpf_rcu_read_lock();
> + n = mapval->node;
> + if (!n) {
> + bpf_rcu_read_unlock();
> + return 2;
> + }
> + bpf_rcu_read_unlock();
> +
> + m = bpf_refcount_acquire(n);
> + if (m)
> + bpf_obj_drop(m);
> +
> + return 0;
> +}
> +
> SEC("?tc")
> __failure __msg("Unreleased reference id=3 alloc_insn={{[0-9]+}}")
> long rbtree_refcounted_node_ref_escapes_owning_input(void *ctx)
The commit message states:
> The verifier previously allowed pointers used after RCU protection ended
> to reach bpf_refcount_acquire() and, for one object layout, a direct write.
> If the object was freed and reused, these operations could access stale
> memory.
>
> Add tests that keep BPF_PROBE_MEM reads accepted but reject reference
> acquisition and direct writes after RCU protection ends. Cover both tested
> object layouts.
Is the description fully accurate?
The "direct write" was never allowed for any object layout. The
PTR_UNTRUSTED check in check_ptr_to_btf_access() predates the series and
rejects writes through any PTR_UNTRUSTED register for both node_data and
node_refcount_only.
The claim "keep BPF_PROBE_MEM reads accepted ... Cover both tested object
layouts" is only true for node_refcount_only (map_kptr_read_after_rcu_unlock).
The node_data read (graph_map_kptr_read_after_spin_unlock) ends up with
PTR_TO_BTF_ID | MEM_ALLOC | PTR_UNTRUSTED | NON_OWN_REF, a combination
missing from bpf_convert_ctx_accesses() in kernel/bpf/fixups.c, so it
stays a plain BPF_MEM load without probe rewrite.
The genuinely new coverage is the two bpf_refcount_acquire() tests. Would
it be more accurate to scope the description to those plus the
refcount-only read?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/31745514567