From: "Darrick J. Wong" <djwong@kernel.org>
To: Dave Chinner <david@fromorbit.com>
Cc: Shiyang Ruan <ruansy.fnst@fujitsu.com>,
linux-kernel@vger.kernel.org, linux-xfs@vger.kernel.org,
nvdimm@lists.linux.dev, linux-mm@kvack.org,
linux-fsdevel@vger.kernel.org, dan.j.williams@intel.com,
hch@infradead.org, jane.chu@oracle.com
Subject: Re: [PATCH 3/3] mm, pmem, xfs: Introduce MF_MEM_REMOVE for unbind
Date: Tue, 20 Sep 2022 17:58:47 -0700 [thread overview]
Message-ID: <Yyphx31m5fO+OZCI@magnolia> (raw)
In-Reply-To: <20220920024519.GQ3600936@dread.disaster.area>
On Tue, Sep 20, 2022 at 12:45:19PM +1000, Dave Chinner wrote:
> On Fri, Sep 02, 2022 at 10:36:01AM +0000, Shiyang Ruan wrote:
> > This patch is inspired by Dan's "mm, dax, pmem: Introduce
> > dev_pagemap_failure()"[1]. With the help of dax_holder and
> > ->notify_failure() mechanism, the pmem driver is able to ask filesystem
> > (or mapped device) on it to unmap all files in use and notify processes
> > who are using those files.
> >
> > Call trace:
> > trigger unbind
> > -> unbind_store()
> > -> ... (skip)
> > -> devres_release_all() # was pmem driver ->remove() in v1
> > -> kill_dax()
> > -> dax_holder_notify_failure(dax_dev, 0, U64_MAX, MF_MEM_PRE_REMOVE)
> > -> xfs_dax_notify_failure()
> >
> > Introduce MF_MEM_PRE_REMOVE to let filesystem know this is a remove
> > event. So do not shutdown filesystem directly if something not
> > supported, or if failure range includes metadata area. Make sure all
> > files and processes are handled correctly.
> >
> > [1]: https://lore.kernel.org/linux-mm/161604050314.1463742.14151665140035795571.stgit@dwillia2-desk3.amr.corp.intel.com/
> >
> > Signed-off-by: Shiyang Ruan <ruansy.fnst@fujitsu.com>
> > ---
> > drivers/dax/super.c | 3 ++-
> > fs/xfs/xfs_notify_failure.c | 23 +++++++++++++++++++++++
> > include/linux/mm.h | 1 +
> > 3 files changed, 26 insertions(+), 1 deletion(-)
> >
> > diff --git a/drivers/dax/super.c b/drivers/dax/super.c
> > index 9b5e2a5eb0ae..cf9a64563fbe 100644
> > --- a/drivers/dax/super.c
> > +++ b/drivers/dax/super.c
> > @@ -323,7 +323,8 @@ void kill_dax(struct dax_device *dax_dev)
> > return;
> >
> > if (dax_dev->holder_data != NULL)
> > - dax_holder_notify_failure(dax_dev, 0, U64_MAX, 0);
> > + dax_holder_notify_failure(dax_dev, 0, U64_MAX,
> > + MF_MEM_PRE_REMOVE);
> >
> > clear_bit(DAXDEV_ALIVE, &dax_dev->flags);
> > synchronize_srcu(&dax_srcu);
> > diff --git a/fs/xfs/xfs_notify_failure.c b/fs/xfs/xfs_notify_failure.c
> > index 3830f908e215..5e04ba7fa403 100644
> > --- a/fs/xfs/xfs_notify_failure.c
> > +++ b/fs/xfs/xfs_notify_failure.c
> > @@ -22,6 +22,7 @@
> >
> > #include <linux/mm.h>
> > #include <linux/dax.h>
> > +#include <linux/fs.h>
> >
> > struct xfs_failure_info {
> > xfs_agblock_t startblock;
> > @@ -77,6 +78,9 @@ xfs_dax_failure_fn(
> >
> > if (XFS_RMAP_NON_INODE_OWNER(rec->rm_owner) ||
> > (rec->rm_flags & (XFS_RMAP_ATTR_FORK | XFS_RMAP_BMBT_BLOCK))) {
> > + /* The device is about to be removed. Not a really failure. */
> > + if (notify->mf_flags & MF_MEM_PRE_REMOVE)
> > + return 0;
> > notify->want_shutdown = true;
> > return 0;
> > }
> > @@ -182,12 +186,23 @@ xfs_dax_notify_failure(
> > struct xfs_mount *mp = dax_holder(dax_dev);
> > u64 ddev_start;
> > u64 ddev_end;
> > + int error;
> >
> > if (!(mp->m_super->s_flags & SB_BORN)) {
> > xfs_warn(mp, "filesystem is not ready for notify_failure()!");
> > return -EIO;
> > }
> >
> > + if (mf_flags & MF_MEM_PRE_REMOVE) {
> > + xfs_info(mp, "device is about to be removed!");
> > + down_write(&mp->m_super->s_umount);
> > + error = sync_filesystem(mp->m_super);
> > + drop_pagecache_sb(mp->m_super, NULL);
> > + up_write(&mp->m_super->s_umount);
> > + if (error)
> > + return error;
>
> If the device is about to go away unexpectedly, shouldn't this shut
> down the filesystem after syncing it here? If the filesystem has
> been shut down, then everything will fail before removal finally
> triggers, and the act of unmounting the filesystem post device
> removal will clean up the page cache and all the other caches.
IIRC they want to kill all the processes with MAP_SYNC mappings sooner
than whenever the admin gets around to unmounting the filesystem, which
is why PRE_REMOVE will then go walk the rmapbt to find processes to
shoot down. I'm not sure, though, if drop_pagecache_sb only touches
DRAM page cache or if it'll shoot down fsdax mappings too?
> IOWs, I don't understand why the page cache is considered special
> here (as opposed to, say, the inode or dentry caches), nor why we
> aren't shutting down the filesystem directly after syncing it to
> disk to ensure that we don't end up with applications losing data as
> a result of racing with the removal....
But yeah, we might as well shut down the fs at the end of PRE_REMOVE
handling, if the rmap walk hasn't already done that.
--D
> Cheers,
>
> Dave.
> --
> Dave Chinner
> david@fromorbit.com
next prev parent reply other threads:[~2022-09-21 0:58 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2022-08-26 13:03 [PATCH v7] mm, pmem, xfs: Introduce MF_MEM_REMOVE for unbind Shiyang Ruan
2022-08-26 21:35 ` Dan Williams
2022-08-29 10:02 ` Shiyang Ruan
2022-08-29 14:49 ` Darrick J. Wong
2022-09-02 10:35 ` [PATCH v8 0/3] " Shiyang Ruan
2022-09-02 10:35 ` [PATCH 1/3] xfs: fix the calculation of length and end Shiyang Ruan
2022-09-14 18:13 ` Darrick J. Wong
2022-09-02 10:36 ` [PATCH 2/3] fs: move drop_pagecache_sb() for others to use Shiyang Ruan
2022-09-14 18:16 ` Darrick J. Wong
2022-09-02 10:36 ` [PATCH 3/3] mm, pmem, xfs: Introduce MF_MEM_REMOVE for unbind Shiyang Ruan
2022-09-14 18:15 ` Darrick J. Wong
2022-09-15 1:49 ` Shiyang Ruan
2022-09-20 2:45 ` Dave Chinner
2022-09-21 0:58 ` Darrick J. Wong [this message]
2022-09-07 9:46 ` [PATCH v8 0/3] " Shiyang Ruan
2022-09-14 18:09 ` Darrick J. Wong
2022-09-14 18:15 ` Darrick J. Wong
2022-09-15 2:56 ` Shiyang Ruan
2022-09-16 15:43 ` Darrick J. Wong
-- strict thread matches above, loose matches on Subject: below --
2022-09-25 13:33 [PATCH v9 " Shiyang Ruan
2022-09-25 13:33 ` [PATCH 3/3] " Shiyang Ruan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=Yyphx31m5fO+OZCI@magnolia \
--to=djwong@kernel.org \
--cc=dan.j.williams@intel.com \
--cc=david@fromorbit.com \
--cc=hch@infradead.org \
--cc=jane.chu@oracle.com \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-xfs@vger.kernel.org \
--cc=nvdimm@lists.linux.dev \
--cc=ruansy.fnst@fujitsu.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.