From: "Darrick J. Wong" <djwong@kernel.org>
To: Christoph Hellwig <hch@lst.de>
Cc: Chandan Babu R <chandan.babu@oracle.com>,
Christian Brauner <brauner@kernel.org>,
linux-xfs@vger.kernel.org, linux-fsdevel@vger.kernel.org
Subject: Re: [PATCH 06/12] xfs: factor out a xfs_file_write_zero_eof helper
Date: Tue, 17 Sep 2024 14:14:19 -0700 [thread overview]
Message-ID: <20240917211419.GC182177@frogsfrogsfrogs> (raw)
In-Reply-To: <20240910043949.3481298-7-hch@lst.de>
On Tue, Sep 10, 2024 at 07:39:08AM +0300, Christoph Hellwig wrote:
> Split a helper from xfs_file_write_checks that just deal with the
> post-EOF zeroing to keep the code readable.
>
> Signed-off-by: Christoph Hellwig <hch@lst.de>
> ---
> fs/xfs/xfs_file.c | 133 ++++++++++++++++++++++++++--------------------
> 1 file changed, 75 insertions(+), 58 deletions(-)
>
> diff --git a/fs/xfs/xfs_file.c b/fs/xfs/xfs_file.c
> index 0d258c21b9897f..a30fda1985e6af 100644
> --- a/fs/xfs/xfs_file.c
> +++ b/fs/xfs/xfs_file.c
> @@ -347,10 +347,70 @@ xfs_file_splice_read(
> return ret;
> }
>
> +static ssize_t
> +xfs_file_write_zero_eof(
> + struct kiocb *iocb,
> + struct iov_iter *from,
> + unsigned int *iolock,
> + size_t count,
> + bool *drained_dio)
> +{
> + struct xfs_inode *ip = XFS_I(iocb->ki_filp->f_mapping->host);
> + loff_t isize;
> +
> + /*
> + * We need to serialise against EOF updates that occur in IO completions
> + * here. We want to make sure that nobody is changing the size while
> + * we do this check until we have placed an IO barrier (i.e. hold
> + * XFS_IOLOCK_EXCL) that prevents new IO from being dispatched. The
> + * spinlock effectively forms a memory barrier once we have
> + * XFS_IOLOCK_EXCL so we are guaranteed to see the latest EOF value and
> + * hence be able to correctly determine if we need to run zeroing.
> + */
> + spin_lock(&ip->i_flags_lock);
> + isize = i_size_read(VFS_I(ip));
> + if (iocb->ki_pos <= isize) {
> + spin_unlock(&ip->i_flags_lock);
> + return 0;
> + }
> + spin_unlock(&ip->i_flags_lock);
> +
> + if (iocb->ki_flags & IOCB_NOWAIT)
> + return -EAGAIN;
> +
> + if (!*drained_dio) {
> + /*
> + * If zeroing is needed and we are currently holding the iolock
> + * shared, we need to update it to exclusive which implies
> + * having to redo all checks before.
> + */
> + if (*iolock == XFS_IOLOCK_SHARED) {
> + xfs_iunlock(ip, *iolock);
> + *iolock = XFS_IOLOCK_EXCL;
> + xfs_ilock(ip, *iolock);
> + iov_iter_reexpand(from, count);
> + }
> +
> + /*
> + * We now have an IO submission barrier in place, but AIO can do
> + * EOF updates during IO completion and hence we now need to
> + * wait for all of them to drain. Non-AIO DIO will have drained
> + * before we are given the XFS_IOLOCK_EXCL, and so for most
> + * cases this wait is a no-op.
> + */
> + inode_dio_wait(VFS_I(ip));
> + *drained_dio = true;
> + return 1;
I gotta say, I'm not a big fan of the "return 1 to loop again" behavior.
Can you add a comment at the top stating that this is a possible return
value and why it gets returned?
--D
> + }
> +
> + trace_xfs_zero_eof(ip, isize, iocb->ki_pos - isize);
> + return xfs_zero_range(ip, isize, iocb->ki_pos - isize, NULL);
> +}
> +
> /*
> * Common pre-write limit and setup checks.
> *
> - * Called with the iolocked held either shared and exclusive according to
> + * Called with the iolock held either shared and exclusive according to
> * @iolock, and returns with it held. Might upgrade the iolock to exclusive
> * if called for a direct write beyond i_size.
> */
> @@ -360,13 +420,10 @@ xfs_file_write_checks(
> struct iov_iter *from,
> unsigned int *iolock)
> {
> - struct file *file = iocb->ki_filp;
> - struct inode *inode = file->f_mapping->host;
> - struct xfs_inode *ip = XFS_I(inode);
> - ssize_t error = 0;
> + struct inode *inode = iocb->ki_filp->f_mapping->host;
> size_t count = iov_iter_count(from);
> bool drained_dio = false;
> - loff_t isize;
> + ssize_t error;
>
> restart:
> error = generic_write_checks(iocb, from);
> @@ -389,7 +446,7 @@ xfs_file_write_checks(
> * exclusively.
> */
> if (*iolock == XFS_IOLOCK_SHARED && !IS_NOSEC(inode)) {
> - xfs_iunlock(ip, *iolock);
> + xfs_iunlock(XFS_I(inode), *iolock);
> *iolock = XFS_IOLOCK_EXCL;
> error = xfs_ilock_iocb(iocb, *iolock);
> if (error) {
> @@ -400,64 +457,24 @@ xfs_file_write_checks(
> }
>
> /*
> - * If the offset is beyond the size of the file, we need to zero any
> + * If the offset is beyond the size of the file, we need to zero all
> * blocks that fall between the existing EOF and the start of this
> - * write. If zeroing is needed and we are currently holding the iolock
> - * shared, we need to update it to exclusive which implies having to
> - * redo all checks before.
> - *
> - * We need to serialise against EOF updates that occur in IO completions
> - * here. We want to make sure that nobody is changing the size while we
> - * do this check until we have placed an IO barrier (i.e. hold the
> - * XFS_IOLOCK_EXCL) that prevents new IO from being dispatched. The
> - * spinlock effectively forms a memory barrier once we have the
> - * XFS_IOLOCK_EXCL so we are guaranteed to see the latest EOF value and
> - * hence be able to correctly determine if we need to run zeroing.
> + * write.
> *
> - * We can do an unlocked check here safely as IO completion can only
> - * extend EOF. Truncate is locked out at this point, so the EOF can
> - * not move backwards, only forwards. Hence we only need to take the
> - * slow path and spin locks when we are at or beyond the current EOF.
> + * We can do an unlocked check for i_size here safely as I/O completion
> + * can only extend EOF. Truncate is locked out at this point, so the
> + * EOF can not move backwards, only forwards. Hence we only need to take
> + * the slow path when we are at or beyond the current EOF.
> */
> - if (iocb->ki_pos <= i_size_read(inode))
> - goto out;
> -
> - spin_lock(&ip->i_flags_lock);
> - isize = i_size_read(inode);
> - if (iocb->ki_pos > isize) {
> - spin_unlock(&ip->i_flags_lock);
> -
> - if (iocb->ki_flags & IOCB_NOWAIT)
> - return -EAGAIN;
> -
> - if (!drained_dio) {
> - if (*iolock == XFS_IOLOCK_SHARED) {
> - xfs_iunlock(ip, *iolock);
> - *iolock = XFS_IOLOCK_EXCL;
> - xfs_ilock(ip, *iolock);
> - iov_iter_reexpand(from, count);
> - }
> - /*
> - * We now have an IO submission barrier in place, but
> - * AIO can do EOF updates during IO completion and hence
> - * we now need to wait for all of them to drain. Non-AIO
> - * DIO will have drained before we are given the
> - * XFS_IOLOCK_EXCL, and so for most cases this wait is a
> - * no-op.
> - */
> - inode_dio_wait(inode);
> - drained_dio = true;
> + if (iocb->ki_pos > i_size_read(inode)) {
> + error = xfs_file_write_zero_eof(iocb, from, iolock, count,
> + &drained_dio);
> + if (error == 1)
> goto restart;
> - }
> -
> - trace_xfs_zero_eof(ip, isize, iocb->ki_pos - isize);
> - error = xfs_zero_range(ip, isize, iocb->ki_pos - isize, NULL);
> if (error)
> return error;
> - } else
> - spin_unlock(&ip->i_flags_lock);
> + }
>
> -out:
> return kiocb_modified(iocb);
> }
>
> --
> 2.45.2
>
>
next prev parent reply other threads:[~2024-09-17 21:14 UTC|newest]
Thread overview: 28+ messages / expand[flat|nested] mbox.gz Atom feed top
2024-09-10 4:39 fix stale delalloc punching for COW I/O v2 Christoph Hellwig
2024-09-10 4:39 ` [PATCH 01/12] iomap: handle a post-direct I/O invalidate race in iomap_write_delalloc_release Christoph Hellwig
2024-09-10 4:39 ` [PATCH 02/12] iomap: improve shared block detection in iomap_unshare_iter Christoph Hellwig
2024-09-10 4:39 ` [PATCH 03/12] iomap: pass flags to iomap_file_buffered_write_punch_delalloc Christoph Hellwig
2024-09-10 4:39 ` [PATCH 04/12] iomap: pass the iomap to the punch callback Christoph Hellwig
2024-09-10 4:39 ` [PATCH 05/12] iomap: remove the iomap_file_buffered_write_punch_delalloc return value Christoph Hellwig
2024-09-10 4:39 ` [PATCH 06/12] xfs: factor out a xfs_file_write_zero_eof helper Christoph Hellwig
2024-09-17 21:14 ` Darrick J. Wong [this message]
2024-09-18 5:09 ` Christoph Hellwig
2024-09-18 15:30 ` Darrick J. Wong
2024-09-20 0:25 ` Dave Chinner
2024-09-20 11:31 ` Christoph Hellwig
2024-09-10 4:39 ` [PATCH 07/12] xfs: take XFS_MMAPLOCK_EXCL xfs_file_write_zero_eof Christoph Hellwig
2024-09-17 21:24 ` Darrick J. Wong
2024-09-18 5:10 ` Christoph Hellwig
2024-09-10 4:39 ` [PATCH 08/12] iomap: zeroing already holds invalidate_lock Christoph Hellwig
2024-09-17 21:29 ` Darrick J. Wong
2024-09-18 5:15 ` Christoph Hellwig
2024-09-18 15:32 ` Darrick J. Wong
2024-09-20 0:42 ` Dave Chinner
2024-09-10 4:39 ` [PATCH 09/12] xfs: support the COW fork in xfs_bmap_punch_delalloc_range Christoph Hellwig
2024-09-10 4:39 ` [PATCH 10/12] xfs: share more code in xfs_buffered_write_iomap_begin Christoph Hellwig
2024-09-10 4:39 ` [PATCH 11/12] xfs: set IOMAP_F_SHARED for all COW fork allocations Christoph Hellwig
2024-09-10 4:39 ` [PATCH 12/12] xfs: punch delalloc extents from the COW fork for COW writes Christoph Hellwig
2024-09-10 9:21 ` fix stale delalloc punching for COW I/O v2 Christian Brauner
2024-09-10 15:17 ` Christoph Hellwig
2024-09-10 9:22 ` (subset) " Christian Brauner
2024-09-10 15:17 ` Christoph Hellwig
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20240917211419.GC182177@frogsfrogsfrogs \
--to=djwong@kernel.org \
--cc=brauner@kernel.org \
--cc=chandan.babu@oracle.com \
--cc=hch@lst.de \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-xfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.