pcpu_init_value() initializes the per-cpu area of a newly created
[lru_]percpu_hash element.  That area is recycled and still holds the
values of whatever element occupied it before, so when the value comes
from a BPF program (onallcpus == false) the function writes the running
CPU's slot and zeroes the rest.

bpf_percpu_hash_update() always passes onallcpus == true, and that arm
calls pcpu_copy_value(), which used to write every CPU.  That changed in
commit c6936161fd55 ("bpf: Add BPF_F_CPU and BPF_F_ALL_CPUS flags
support for percpu_hash and lru_percpu_hash maps"): with BPF_F_CPU it
writes the one CPU named in map_flags and returns.  On the create
path the remaining slots are left as they were, and a lookup of the new
key hands back the recycled element's values:

  update(k1, 0xdeadc0de, BPF_F_ALL_CPUS)  every CPU holds 0xdeadc0de
  delete(k1)                              element back on the freelist
  update(k2, 0xc0ffee, BPF_F_CPU | 0)     creates, writes CPU 0 only
  lookup(k2)                              CPU 0 0xc0ffee, rest 0xdeadc0de

Commit d3bec0138bfb ("bpf: Zero-fill re-used per-cpu map element")
established that a re-used element must not return the previous
tenant's values.  BPF_F_CPU is the first way to reach pcpu_init_value()
writing a single CPU with onallcpus set, so extend the zero-filling arm
to cover it, with map_flags >> 32 naming the CPU that receives the
value.

Only creation is affected: pcpu_init_value() is reached from the two
create branches, while an update of an existing element goes straight to
pcpu_copy_value(), where writing one CPU and leaving the others is the
point of the flag.

Fixes: c6936161fd55 ("bpf: Add BPF_F_CPU and BPF_F_ALL_CPUS flags support for 
percpu_hash and lru_percpu_hash maps")
Signed-off-by: Donggeun Yoo <[email protected]>
---
Tested on x86_64 under QEMU/KVM against bpf/master a11212910cf0: with
this patch the selftest in 2/2 passes on all three allocation modes,
and without it all three read 0xdeadc0de where they expect 0.  Numbers
in the cover letter.

The merged arm no longer calls bpf_obj_cancel_fields() on the named
CPU.  That call is inert on this path: it acts only on BPF_TIMER,
BPF_WORKQUEUE and BPF_TASK_WORK, and map_check_btf() rejects all three
for [lru_]percpu_hash.

 kernel/bpf/hashtab.c | 6 +++---
 1 file changed, 3 insertions(+), 3 deletions(-)

diff --git a/kernel/bpf/hashtab.c b/kernel/bpf/hashtab.c
index 4f495dcbf670c..c4683d0e4c149 100644
--- a/kernel/bpf/hashtab.c
+++ b/kernel/bpf/hashtab.c
@@ -1056,12 +1056,12 @@ static void pcpu_init_value(struct bpf_htab *htab, void 
__percpu *pptr,
         * known initial values for cpus other than current one
         * (onallcpus=false always when coming from bpf prog).
         */
-       if (!onallcpus) {
-               int current_cpu = raw_smp_processor_id();
+       if (!onallcpus || (map_flags & BPF_F_CPU)) {
+               int init_cpu = onallcpus ? map_flags >> 32 : 
raw_smp_processor_id();
                int cpu;
 
                for_each_possible_cpu(cpu) {
-                       if (cpu == current_cpu)
+                       if (cpu == init_cpu)
                                copy_map_value(&htab->map, per_cpu_ptr(pptr, 
cpu), value);
                        else /* Since elem is preallocated, we cannot touch 
special fields */
                                zero_map_value(&htab->map, per_cpu_ptr(pptr, 
cpu));
-- 
2.53.0


Reply via email to