All of lore.kernel.org
 help / color / mirror / Atom feed
From: Al Viro <viro@ZenIV.linux.org.uk>
To: Omar Sandoval <osandov@osandov.com>
Cc: linux-fsdevel@vger.kernel.org, kernel-team@fb.com,
	"Eric W . Biederman" <ebiederm@xmission.com>,
	Tejun Heo <tj@kernel.org>
Subject: Re: [PATCH] fs: only sync() superblocks reachable from the current namespace
Date: Fri, 26 Jan 2018 23:13:51 +0000	[thread overview]
Message-ID: <20180126231350.GC13338@ZenIV.linux.org.uk> (raw)
In-Reply-To: <05434cda5cc3b461b5d70467b094904ad23fdc11.1517007510.git.osandov@fb.com>

On Fri, Jan 26, 2018 at 02:58:39PM -0800, Omar Sandoval wrote:
> From: Omar Sandoval <osandov@fb.com>
> 
> Currently, the sync() syscall is system-wide, so any process in a
> container can cause significant I/O stalls across the system by calling
> sync(). This is even true for filesystems which are not accessible in
> the process' mount namespace. This patch scopes sync() to only write out
> filesystems reachable in the current mount namespace, except for the
> initial mount namespace, which still syncs everything to avoid
> surprises. This fixes the broken isolation we were seeing here.

> +static int sb_reachable(struct super_block *sb, struct mnt_namespace *mnt_ns)
> +{
> +	struct mount *mnt;
> +
> +	if (!mnt_ns)
> +		return 1;
> +
> +	list_for_each_entry(mnt, &sb->s_mounts, mnt_instance) {
> +		if (mnt->mnt_ns == mnt_ns)
> +			return 1;
> +	}
> +	return 0;
> +}

Erm...  And just what is protecting the list here?

>  static void fdatawrite_one_bdev(struct block_device *bdev, void *arg)
> @@ -107,12 +138,18 @@ static void fdatawait_one_bdev(struct block_device *bdev, void *arg)
>   */
>  SYSCALL_DEFINE0(sync)
>  {
> -	int nowait = 0, wait = 1;
> +	struct sb_sync arg = {
> +		.mnt_ns = current->nsproxy->mnt_ns,
> +	};
> +
> +	if (arg.mnt_ns == init_task.nsproxy->mnt_ns)
> +		arg.mnt_ns = NULL;
>  
>  	wakeup_flusher_threads(WB_REASON_SYNC);
> -	iterate_supers(sync_inodes_one_sb, NULL);
> -	iterate_supers(sync_fs_one_sb, &nowait);
> -	iterate_supers(sync_fs_one_sb, &wait);
> +	iterate_supers(sync_inodes_one_sb, &arg);
> +	iterate_supers(sync_fs_one_sb, &arg);
> +	arg.wait = 1;
> +	iterate_supers(sync_fs_one_sb, &arg);

So now sync() includes O(total vfsmounts on the system) walking the lists, no
matter what *and* in a situation when a lazy-unmounted filesystem is held active
by an opened file sync(2) won't touch that filesystem.  Unless done in the
magical namespace init(8) happens to run in.

  parent reply	other threads:[~2018-01-26 23:13 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2018-01-26 22:58 [PATCH] fs: only sync() superblocks reachable from the current namespace Omar Sandoval
2018-01-26 23:00 ` Omar Sandoval
2018-01-26 23:13 ` Al Viro [this message]
2018-01-26 23:17   ` Al Viro
2018-01-26 23:46     ` Omar Sandoval

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20180126231350.GC13338@ZenIV.linux.org.uk \
    --to=viro@zeniv.linux.org.uk \
    --cc=ebiederm@xmission.com \
    --cc=kernel-team@fb.com \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=osandov@osandov.com \
    --cc=tj@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.