All of lore.kernel.org
 help / color / mirror / Atom feed
From: Giuseppe Scrivano <gscrivan@redhat.com>
To: Christian Brauner <brauner@kernel.org>
Cc: linux-erofs@lists.ozlabs.org,  xiang@kernel.org,
	linux-fsdevel@vger.kernel.org,  amir73il@gmail.com
Subject: Re: [PATCH v4] erofs: reuse superblock for file-backed mounts
Date: Mon, 10 Aug 2026 17:11:46 +0200	[thread overview]
Message-ID: <874ih2rowd.fsf@redhat.com> (raw)
In-Reply-To: <20260810-losging-beengt-zarte-1047dd62c195@brauner> (Christian Brauner's message of "Mon, 10 Aug 2026 15:34:19 +0200")

Hi Christian,

Christian Brauner <brauner@kernel.org> writes:

> Hi Giuseppe,
>
>> When the same file-backed image is mounted multiple times (via path or
>> fd) with the superblock_share mount option, reuse the existing
>> superblock instead of creating a new one.  This allows multiple mounts
>> of the same image to share in-kernel data structures more efficiently.
>> 
>> The backing file is identified by its inode and fsoffset.  If mount
>> options conflict or extra devices are used, a separate superblock
>> is created transparently as a fallback.  Remount is not allowed on
>> shared superblocks.
>> 
>> Tested by mounting a 177M Fedora EROFS image 20 times with full
>> traversal:
>> 
>>   #!/bin/sh
>>   IMG=${1:-/root/fedora.erofs}
>>   OPTS=${2:+-o $2}
>>   N=20
>>   DIR=$(mktemp -d)
>>   trap "umount $DIR/m* 2>/dev/null; rm -rf $DIR" EXIT
>> 
>>   sync; echo 3 > /proc/sys/vm/drop_caches
>>   INODES_BEFORE=$(grep erofs_inode /proc/slabinfo | awk '{print $2}')
>>   MEM_BEFORE=$(grep ^Slab: /proc/meminfo | awk '{print $2}')
>> 
>>   for i in $(seq 1 $N); do
>>       mkdir $DIR/m$i
>>       mount -t erofs $OPTS "$IMG" $DIR/m$i
>>       find $DIR/m$i > /dev/null
>>   done
>> 
>>   echo "Superblocks: $(grep $DIR /proc/self/mountinfo | \
>>       awk '{print $3}' | sort -u | wc -l)"
>>   echo "erofs_inode delta: +$(( $(grep erofs_inode /proc/slabinfo | \
>>       awk '{print $2}') - INODES_BEFORE ))"
>>   echo "Slab delta: +$(( $(grep ^Slab: /proc/meminfo | \
>>       awk '{print $2}') - MEM_BEFORE )) kB"
>> 
>> without superblock_share:
>> 
>>   # time ./test.sh /root/fedora.erofs
>>   Superblocks: 20
>>   erofs_inode delta: +45864
>>   Slab delta: +37448 kB
>>   real	0m2.311s
>>   user	0m0.222s
>>   sys	0m1.998s
>> 
>> with superblock_share:
>> 
>>   # time ./test.sh /root/fedora.erofs superblock_share
>>   Superblocks: 1
>>   erofs_inode delta: +2044
>>   Slab delta: +284 kB
>>   real	0m0.679s
>>   user	0m0.157s
>>   sys	0m0.469s
>> 
>> The time difference shows that sharing the superblock also benefits
>> the page cache and inode cache, as subsequent mounts of the same image
>> avoid re-reading the backing file.  This is particularly useful for
>> container hosts running multiple containers from the same base image.
>> 
>> Signed-off-by: Giuseppe Scrivano <gscrivan@redhat.com>
>>
>> diff --git a/Documentation/filesystems/erofs.rst b/Documentation/filesystems/erofs.rst
>> index d301d9ac946a..49c4e7dbc5d6 100644
>> --- a/Documentation/filesystems/erofs.rst
>> +++ b/Documentation/filesystems/erofs.rst
>> @@ -139,6 +139,9 @@ inode_share            Enable inode page sharing for this filesystem.  Inodes wi
>>                         page cache.
>>  source=%s              (For file-backed mounts) Specify the backing image as a path
>>                         or as an already-opened file descriptor.
>> +superblock_share       (For file-backed mounts) Share the superblock when the same
>> +                       backing file is mounted more than once with compatible
>> +                       options.  Remount is not allowed on shared superblocks.
>>  ===================    =========================================================
>>  
>>  File-backed mounts
>> @@ -156,6 +159,11 @@ Only regular files are accepted as backing files; to mount an image that
>>  resides on a block device, use the traditional block device mount path
>>  instead.
>>  
>> +When superblock_share is specified and the same backing file (identified
>> +by inode and fsoffset) is mounted more than once with compatible mount
>> +options, the kernel reuses the existing superblock instead of creating a
>> +new one.  Remount is not allowed on shared superblocks.
>> +
>>  Sysfs Entries
>>  =============
>>  
>> diff --git a/fs/erofs/internal.h b/fs/erofs/internal.h
>> index 57bd21859c65..86da219f8f4e 100644
>> --- a/fs/erofs/internal.h
>> +++ b/fs/erofs/internal.h
>> @@ -156,6 +156,7 @@ struct erofs_sb_info {
>>  #define EROFS_MOUNT_DAX_NEVER		0x00000080
>>  #define EROFS_MOUNT_DIRECT_IO		0x00000100
>>  #define EROFS_MOUNT_INODE_SHARE		0x00000200
>> +#define EROFS_MOUNT_SUPERBLOCK_SHARE	0x00000400
>>  
>>  #define clear_opt(opt, option)	((opt)->mount_opt &= ~EROFS_MOUNT_##option)
>>  #define set_opt(opt, option)	((opt)->mount_opt |= EROFS_MOUNT_##option)
>> diff --git a/fs/erofs/super.c b/fs/erofs/super.c
>> index 3d92caec8d3a..4e926376e17b 100644
>> --- a/fs/erofs/super.c
>> +++ b/fs/erofs/super.c
>> @@ -386,7 +386,7 @@ static void erofs_default_options(struct erofs_sb_info *sbi)
>>  enum {
>>  	Opt_user_xattr, Opt_acl, Opt_cache_strategy, Opt_dax, Opt_dax_enum,
>>  	Opt_device, Opt_domain_id, Opt_directio, Opt_fsoffset, Opt_inode_share,
>> -	Opt_source,
>> +	Opt_source, Opt_superblock_share,
>>  };
>>  
>>  static const struct constant_table erofs_param_cache_strategy[] = {
>> @@ -415,6 +415,7 @@ static const struct fs_parameter_spec erofs_fs_parameters[] = {
>>  	fsparam_u64("fsoffset",			Opt_fsoffset),
>>  	fsparam_flag("inode_share",		Opt_inode_share),
>>  	fsparam_file_or_string("source",	Opt_source),
>> +	fsparam_flag("superblock_share",	Opt_superblock_share),
>>  	{}
>>  };
>>  
>> @@ -560,6 +561,12 @@ static int erofs_fc_parse_param(struct fs_context *fc,
>>  		else
>>  			set_opt(&sbi->opt, INODE_SHARE);
>>  		break;
>> +	case Opt_superblock_share:
>> +		if (!IS_ENABLED(CONFIG_EROFS_FS_BACKED_BY_FILE))
>> +			errorfc(fc, "%s option not supported", erofs_fs_parameters[opt].name);
>> +		else
>> +			set_opt(&sbi->opt, SUPERBLOCK_SHARE);
>> +		break;
>>  	case Opt_source:
>>  		return erofs_fc_parse_source(fc, param);
>>  	}
>> @@ -652,6 +659,17 @@ static int erofs_fc_fill_super(struct super_block *sb, struct fs_context *fc)
>>  		errorfc(fc, "FSDAX is not allowed when inode_share is on");
>>  		return -EINVAL;
>>  	}
>> +	if (test_opt(&sbi->opt, SUPERBLOCK_SHARE) &&
>> +	    test_opt(&sbi->opt, INODE_SHARE)) {
>> +		errorfc(fc, "superblock_share is not allowed when inode_share is on");
>> +		return -EINVAL;
>> +	}
>> +	/* Extra devices are not yet supported with superblock sharing.  */
>> +	if (test_opt(&sbi->opt, SUPERBLOCK_SHARE) &&
>> +	    sbi->devs->extra_devices) {
>> +		errorfc(fc, "superblock_share does not support extra devices");
>> +		return -EINVAL;
>> +	}
>>  
>>  	sbi->blkszbits = PAGE_SHIFT;
>>  	if (!sb->s_bdev) {
>> @@ -779,6 +797,51 @@ static int erofs_fc_fill_super(struct super_block *sb, struct fs_context *fc)
>>  	return 0;
>>  }
>>  
>> +static int erofs_fc_test_file_super(struct super_block *sb,
>> +				    struct fs_context *fc)
>> +{
>> +	struct erofs_sb_info *sbi = EROFS_SB(sb);
>> +	struct erofs_sb_info *new_sbi = fc->s_fs_info;
>> +
>> +	if (sb->s_iflags & SB_I_RETIRED)
>> +		return 0;
>> +	if (!sbi->dif0.file || !new_sbi->dif0.file)
>> +		return 0;
>> +	if (!test_opt(&new_sbi->opt, SUPERBLOCK_SHARE))
>> +		return 0;
>> +	return file_inode(sbi->dif0.file) == file_inode(new_sbi->dif0.file) &&
>> +	       sbi->dif0.fsoff == new_sbi->dif0.fsoff &&
>> +	       !memcmp(&sbi->opt, &new_sbi->opt, sizeof(sbi->opt));
>> +}
>> +
>> +static int erofs_fc_get_tree_file(struct fs_context *fc)
>> +{
>> +	struct erofs_sb_info *sbi = fc->s_fs_info;
>> +	struct super_block *sb;
>> +	int err;
>> +
>> +	if (!test_opt(&sbi->opt, SUPERBLOCK_SHARE))
>> +		return get_tree_nodev(fc, erofs_fc_fill_super);
>> +
>> +	sb = sget_fc(fc, erofs_fc_test_file_super, set_anon_super_fc);
>> +	if (IS_ERR(sb))
>> +		return PTR_ERR(sb);
>> +
>> +	if (sb->s_root) {
>> +		erofs_info(sb, "sharing superblock for the same backing file");
>> +	} else {
>> +		err = erofs_fc_fill_super(sb, fc);
>> +		if (err) {
>> +			deactivate_locked_super(sb);
>> +			return err;
>> +		}
>> +		sb->s_flags |= SB_ACTIVE;
>> +	}
>> +
>> +	fc->root = dget(sb->s_root);
>> +	return 0;
>> +}
>
> I think you should rename vfs_get_super() to get_tree_super() and expose
> it to modules. Then this function collapses to:
>
> if (!test_opt(&sbi->opt, SUPERBLOCK_SHARE))
> 	get_tree_super(fc, NULL, erofs_fc_fill_super);
>
> return get_tree_super(fc, erofs_fc_test_file_super, erofs_fc_fill_super);
>
> or simpler, if you move the sbi test for SUPERBLOCK_SHARE into
> erofs_fc_test_file_super:
>
> get_tree_super(fc, erofs_fc_test_file_super, erofs_fc_fill_super);

that is neat!  I'll send a new version with this fix.

Regards,
Giuseppe


      reply	other threads:[~2026-08-10 15:12 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-04  6:00 [PATCH v4] erofs: reuse superblock for file-backed mounts Giuseppe Scrivano
2026-08-10 13:34 ` Christian Brauner
2026-08-10 15:11   ` Giuseppe Scrivano [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=874ih2rowd.fsf@redhat.com \
    --to=gscrivan@redhat.com \
    --cc=amir73il@gmail.com \
    --cc=brauner@kernel.org \
    --cc=linux-erofs@lists.ozlabs.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=xiang@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.