From: Giuseppe Scrivano <gscrivan@redhat.com>
To: Gao Xiang <hsiangkao@linux.alibaba.com>
Cc: Christian Brauner <brauner@kernel.org>,
linux-erofs@lists.ozlabs.org, cyphar@cyphar.com,
linux-fsdevel@vger.kernel.org
Subject: Re: [PATCH v4] erofs: accept source file descriptor via fsconfig
Date: Mon, 27 Jul 2026 10:01:34 +0200 [thread overview]
Message-ID: <87mrvckgip.fsf@redhat.com> (raw)
In-Reply-To: <95a5fc3a-4259-44a7-bb72-8ff35a49a26f@linux.alibaba.com> (Gao Xiang's message of "Fri, 24 Jul 2026 06:38:20 +0800")
Gao Xiang <hsiangkao@linux.alibaba.com> writes:
> Hi Christian,
>
> On 2026/7/23 22:46, Christian Brauner wrote:
>>> I'm not quite sure if I catched the point, I think Giuseppe's patch here
>>> tried to record `file` into `sbi->dif0.file` (which indicates the primary
>>> "device" later.)
>>>
>>> And if `sbi->dif0.file` is set up by erofs_fc_parse_source(),
>>> erofs_fc_get_tree() will just use `sbi->dif0.file` instead of
>>> `fc->source` according to this patch.
>> Oh, so you only do it for file-backed mounts. Do you only allow
>> regular
>> files or do you also support block devices with
>> CONFIG_EROFS_FS_BACKED_BY_FILE?
>
> Block devices with CONFIG_EROFS_FS_BACKED_BY_FILE are supported,
> but with only `fc->source` (not this way.)
>
> That is the limitation I see in Giuseppe's patch. I'd hoped
> bdev-backed mounts could work the same way, but that would require
> changes to the VFS flow.
>
> Since this is a side improvement, I think it's fine as long as
> it's documented somewhere, and I do hope Giuseppe can at least
> address the documentation.
would something like the following be enough?
diff --git a/Documentation/filesystems/erofs.rst b/Documentation/filesystems/erofs.rst
index 4230884fb359..768e1d43dfcc 100644
--- a/Documentation/filesystems/erofs.rst
+++ b/Documentation/filesystems/erofs.rst
@@ -139,6 +139,29 @@ inode_share Enable inode page sharing for this filesystem. Inodes wi
page cache.
=================== =========================================================
+File-backed mounts
+==================
+
+When ``CONFIG_EROFS_FS_BACKED_BY_FILE`` is enabled, EROFS can mount filesystem
+images stored as regular files directly, without requiring a loopback block
+device. The source can be specified either by path or by passing an
+already-opened file descriptor via ``fsconfig(fd, FSCONFIG_SET_FD, "source",
+NULL, source_fd)``. Only regular files are accepted; block devices must use
+the standard block device mount path.
+
+The backing file content must remain stable for the lifetime of the mount.
+EROFS never writes to it, but concurrent modifications by other processes lead
+to undefined behavior.
+
+Ioctls
+======
+
+``EROFS_IOC_GET_SOURCE_FD``
+ Return a read-only file descriptor (``O_CLOEXEC``) for the backing file of a
+ file-backed mount. Returns ``-ENOENT`` on block-device-backed mounts.
+ Requires ``CAP_SYS_ADMIN`` in the initial user namespace (returns ``-EPERM``
+ otherwise).
+
Sysfs Entries
=============
Regards,
Giuseppe
> I'd also like to make sure the way fc->source is filled
> out of fd passing follows common practice, so that if fd-based
> bdev-backed mounts land in the VFS later, they can keep
> the same fc->source convention, otherwise it will cause
> a userspace behavior change.
>
>> Do you document the expected behavior for the file you're consuming?
>> Meaning, are concurrent modifications supported and what type of
>> behavior does this exhibit?
>
>
> As I perhaps mentioned, EROFS itself (or many EROFS) won't do any
> modification to the underlayfs bdev or files by design so the
> standard behavior is the blob devices / files won't get any change.
> Beyond that, both the on-disk format and the implementation are
> designed to tolerate unexpected external modifications (or storage
> media damage). Even in the worst case, where the underlying storage
> (block device or backing filesystem) is malicious, corrupted on-disk
> (meta)data will not lead to the kind of complex, hard-to-resolve
> inconsistencies you see in general-purpose writable filesystems,
> whose ondisk/in-memory cached metadata is much harder to reconcile.
>
> I'm not sure whether you'll agree, but I want to emphasize that
> again this is one of EROFS core design goals: the on-disk and
> implementation design ensure that. If there is any human bug, it
> will be addressed and fixed as long as it discloses: it won't be
> hard to fixed.
>
> But if you really want to avoid concurrent modifications or keep
> the image golden, I think dmverity or fsverify should be enforced
> to ensure the filesystem won't be modified unexpectedly or expectedly.
>
>>
>>> The reason why `fc->source` is set was discussed in the
>>> thread of the previous version suggested by Aleksa.
>>>
>>>>
>>>> If they close it before this means you can mount something completely
>>>> different. The other thing is even if they keep the fd open someone
>>>> could just rename the damn thing and fc->source ends up pointing
>>>> somwhere completely different. The could switch namespaces as well in
>>>> some circumstances and then it points again into wherever.
>>>>
>>>> I've played with that fd idea before. The only way to make this work
>>>> correctly is if you plumb this down into get_tree_nodev()
>>>
>>> fc->source in this case has no use in erofs_fc_get_tree() (`fc->source`
>>> is just used for mountinfo for example), `sbi->dif0.file` works instead
>>> I hope I don't misunderstand something.
>> No, I misunderstood this.
>
> Thanks,
> Gao Xiang
next prev parent reply other threads:[~2026-07-27 8:01 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-17 13:41 [PATCH v4] erofs: accept source file descriptor via fsconfig Giuseppe Scrivano
2026-07-20 9:45 ` Gao Xiang
2026-07-22 16:11 ` Christian Brauner
2026-07-22 16:25 ` Gao Xiang
2026-07-23 14:46 ` Christian Brauner
2026-07-23 22:38 ` Gao Xiang
2026-07-27 8:01 ` Giuseppe Scrivano [this message]
2026-07-28 12:33 ` Gao Xiang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=87mrvckgip.fsf@redhat.com \
--to=gscrivan@redhat.com \
--cc=brauner@kernel.org \
--cc=cyphar@cyphar.com \
--cc=hsiangkao@linux.alibaba.com \
--cc=linux-erofs@lists.ozlabs.org \
--cc=linux-fsdevel@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).