On 9/8/26 12:34 PM, Mike Snitzer wrote:
> nfsd_dio_iter_is_aligned() approves a write iterator against the
> file's STATX_DIOALIGN attributes (a whole-iterator iov_iter_alignment()
> test against dio_mem_align), but the block stack applies stricter
> geometry tests at bio split time: bio_split_io_at() checks each bvec's
> offset and length against the queue's dma_alignment and may find no
> valid block-size-aligned split at all. An ITER_BVEC WRITE payload can
> pass the former and fail the latter: bio_iov_bvec_set() hands nfsd's
> bvec array to the queue as-is, and the payload's first fragment starts
> mid-page (the RPC header precedes it in the receive buffer), so the
> iterator's interior page boundaries need not be logical-block aligned
> and a bio the queue must split may have no valid split point. When
> that happens, nfsd_direct_write() returned the -EINVAL to the client
> as a failed WRITE (NFS4ERR_INVAL) -- for a perfectly valid request.
>
> Observed against a brd-backed nvme-loop XFS export (dio_mem_align=4)
> with 1 MiB WRITEs, e.g. arriving as 65 bvecs with bv0=(408,15976):
> the gate admits the iterator, the block layer rejects it, and every
> large write on the affected connection errors out (dd: Invalid
> argument).

Thanks for chasing this down. The bv0 numbers make the gate defect
clear: 15976 is not a multiple of the logical block size, so the
direct segment's first interior bvec boundary lands mid-sector. The
boundaries after that are page boundaries, which are fine.

A small correction for the commit message: nfsd_dio_iter_is_aligned()
doesn't exist. The gate is the first-bvec offset test in
nfsd_write_dio_iters_init(), and it checks only that one offset
against nf_dio_mem_align. Likewise bio_iov_bvec_set() is now
bio_iov_iter_set().

> Treat -EINVAL from the direct attempt as "not direct-able": restore the
> segment's iterator and retry it as (uncached when FOP_DONTCACHE)
> buffered I/O, the same fallback nfsd_write_dio_iters_init() picks for
> geometries it rejects itself.

[ ... ]

> @@ -1467,6 +1469,33 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct 
> svc_fh *fhp,
>               expected = iov_iter_count(&segments[i].iter);
>
>               host_err = vfs_iocb_iter_write(file, kiocb, &segments[i].iter);
> +             if (unlikely(host_err == -EINVAL &&
> +                          (kiocb->ki_flags & IOCB_DIRECT))) {

[ ... ]

> +                     segments[i].iter = saved_iter;
> +                     kiocb->ki_flags &= ~IOCB_DIRECT;
> +                     if (file->f_op->fop_flags & FOP_DONTCACHE)
> +                             kiocb->ki_flags |= IOCB_DONTCACHE;
> +                     trace_nfsd_write_vector(rqstp, fhp, kiocb->ki_pos,
> +                                             segments[i].iter.count);
> +                     host_err = vfs_iocb_iter_write(file, kiocb,
> +                                                    &segments[i].iter);
> +             }

Per our discussion last October:

  https://lore.kernel.org/linux-nfs/[email protected]/

The conclusion then was that -EINVAL from ->write_iter can come from
a number of conditions in the filesystem, so NFSD can't treat it as
meaning only that the I/O was misaligned. That still holds, so I'd
rather not use -EINVAL to signal a retry. An -EINVAL that really is
the filesystem rejecting the request would now cost a second full
write attempt before surfacing anyway.

nfsd_write_dio_iters_init() already has the segment start and
nf_dio_offset_align, and after the first bvec every boundary is
page-aligned. If it also requires the first bvec's remaining length
(from the segment start) to be a multiple of offset_align and takes
the no_dio path otherwise, that rejects bv0=(408,15976) up front
using only data NFSD already has.

What would help me understand the failure even better:

- Which -EINVAL in bio_split_io_at() fired: the per-bvec dma_alignment
  test, or the zero-length result after ALIGN_DOWN()?

- On the reproducer, how does stx_dio_offset_align compare with the
  queue's logical_block_size?

If there turn out to be cases the gate can't predict from the statx
data, that seems like a question for the block and fs folks about
what error the filesystem should surface, rather than something to
work around in NFSD.


-- 
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)

Reply via email to