> blk-iocost: add BPF struct_ops cost model support

TL;DR: Possible behaviour issue in blk-iocost BPF model: a blkg whose
radix_tree_insert() fails may get iocg_init() without iocg_free(), or the
reverse, if attach/detach races, contradicting the documented pairing. No
kernel harm.

> diff --git a/block/blk-iocost.c b/block/blk-iocost.c
> index 2745bffcd5ee..581b48543441 100644
> --- a/block/blk-iocost.c
> +++ b/block/blk-iocost.c

[ ... ]

> @@ -3048,6 +3163,18 @@ static void ioc_pd_init(struct blkg_policy_data *pd)
>       spin_lock_irqsave(&ioc->lock, flags);
>       weight_updated(iocg, &now);
>       spin_unlock_irqrestore(&ioc->lock, flags);
> +
> +     /*
> +      * the attached model is RCU-protected: a concurrent detach
> +      * publishes NULL and the struct_ops image survives it by a
> +      * grace period, so the callback is safe inside the read-side
> +      * critical section
> +      */
> +     rcu_read_lock();
> +     model = ioc_attached_or_null(ioc);
> +     if (model && model->iocg_init)
> +             model->iocg_init(blkg->blkcg, ioc->rqos.disk->queue);
> +     rcu_read_unlock();
>  }
>
>  static void iocg_release(struct rcu_head *rcu)
> @@ -3066,8 +3193,15 @@ static void ioc_pd_free(struct blkg_policy_data *pd)
>       struct blkcg_gq *blkg = pd_to_blkg(pd);
>       struct ioc *ioc = iocg->ioc;
>       unsigned long flags;
> +     const struct iocost_model_ops *model;
>
>       if (ioc) {
> +             rcu_read_lock();
> +             model = ioc_attached_or_null(ioc);
> +             if (model && model->iocg_free)
> +                     model->iocg_free(blkg->blkcg, ioc->rqos.disk->queue);
> +             rcu_read_unlock();
> +

The commit message says:

  "... iocg_init() is delivered on attach to every cgroup which already
  has a blkg on the device and to each one appearing afterwards,
  iocg_free() on detach to every cgroup still existing then, and at
  policy deactivation time for the rest, so init and free always pair
  up."

Does that pairing hold for a blkg that never reaches q->blkg_list?

blkg_create() calls the policy's pd_init_fn (ioc_pd_init(), which delivers
iocg_init() when a model is attached) before the radix_tree_insert(), and
only adds the blkg to q->blkg_list if that insert succeeds:

blkg_create() {
    ...
    pol->pd_init_fn(blkg->pd[i]);
    ...
    ret = radix_tree_insert(&blkcg->blkg_tree, disk->queue->id, blkg);
    if (likely(!ret)) {
        ...
        list_add(&blkg->q_node, &disk->queue->blkg_list);
    ...
    /* @blkg failed fully initialized, use the usual release path */
    percpu_ref_kill(&blkg->refcnt);
    return ERR_PTR(ret);
}

The insert is into a GFP_NOWAIT radix tree, and on the bio association path
(blkg_tryget_closest() -> blkg_lookup_create() -> blkg_create()) there is no
preload, so it can fail with -ENOMEM. That blkg is then released through
blkg_release() -> call_rcu() -> blkg_free_workfn() -> ioc_pd_free().

If ioc_bpf_detach() runs after that ioc_pd_init() but before that
ioc_pd_free(), ioc_bpf_walk_iocgs() does not find the blkg on q->blkg_list,
and ioc_pd_free() then sees ioc->attached as NULL. iocg_init() has been
delivered and iocg_free() never is.

The reverse looks possible too.  If ioc_bpf_attach() runs in that window,
ioc_pd_init() ran with no model, the attach walk misses the blkg, and
ioc_pd_free() then delivers an iocg_free() with no matching iocg_init().

Can a BPF model that keeps per-cgroup state in iocg_init()/iocg_free() see
these unpaired callbacks?  Nothing in the kernel itself is harmed, but the
commit message and the comment above struct iocost_model_ops say the
callbacks always pair up.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37087498848

Reply via email to