From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from relay.hostedemail.com (smtprelay0012.hostedemail.com [216.40.44.12]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5777442BEAF; Thu, 6 Aug 2026 22:23:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=216.40.44.12 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786054987; cv=none; b=gvypwkZ20Dd7p1oAZOTv4TfXPwLhWtWY4CK6DE+5YTHvjO8xDlWpWai/3YKixfcEp+/6dbfV1yoLkiXNh4+FQQVNUCwC1bZfT4p1Mj94nY93YmuhuMMJ3ouWGu2dnZVwQxKRdif3y5HWljhi7yBhn7lEG0GFAtt7pDM03nMNtZ8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786054987; c=relaxed/simple; bh=tGvzdTq7uHgUkxF6mM6FTnwURgWycNfryP9u0aplqSQ=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=Vo9MjLpLekJWc3rHAkzIoGtV7gRge2tPJL3Tqv+RxdvzvjvwllYhbFVSPS4h4Yckr2mSuO2bRL215G6TrfqGV+cwlMh4LaLAFhkIztE1iOX1ospyZKXCsqHKO4ygo1HNPYRbgDs5IOolZsEg0I8cVJvhubGDVUFEizP380VSLDY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=groves.net; spf=pass smtp.mailfrom=groves.net; arc=none smtp.client-ip=216.40.44.12 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=groves.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=groves.net Received: from omf04.hostedemail.com (lb01a-stub [10.200.18.249]) by unirelay03.hostedemail.com (Postfix) with ESMTP id 35084A01FA; Thu, 6 Aug 2026 22:23:04 +0000 (UTC) Received: from [HIDDEN] (Authenticated sender: john@groves.net) by omf04.hostedemail.com (Postfix) with ESMTPA id 5165220024; Thu, 6 Aug 2026 22:22:47 +0000 (UTC) Date: Thu, 6 Aug 2026 17:22:46 -0500 From: John Groves To: "Darrick J. Wong" Cc: John Groves , Miklos Szeredi , Dan Williams , Bernd Schubert , Alison Schofield , John Groves , Jonathan Corbet , Jake Edge , Shuah Khan , Vishal Verma , Dave Jiang , Matthew Wilcox , Jan Kara , Alexander Viro , David Hildenbrand , Christian Brauner , Randy Dunlap , Jeff Layton , Amir Goldstein , Jonathan Cameron , Stefan Hajnoczi , Joanne Koong , Josef Bacik , Bagas Sanjaya , Chen Linxuan , James Morse , Fuad Tabba , Sean Christopherson , Shivank Garg , Ackerley Tng , Gregory Price , Andrew Morton , Namjae Jeon , Lorenzo Stoakes , Greg Kroah-Hartman , Ira Weiny , Pasha Tatashin , Haren Myneni , Pratyush Yadav , Giovanni Cabiddu , Jiri Slaby , Ethan Nelson-Moore , Gabriel Whigham , Aravind Ramesh , Ajay Joshi , "venkataravis@micron.com" , "linux-doc@vger.kernel.org" , "linux-kernel@vger.kernel.org" , "nvdimm@lists.linux.dev" , "linux-cxl@vger.kernel.org" , "linux-fsdevel@vger.kernel.org" , "fuse-devel@lists.linux.dev" Subject: Re: [PATCH V12 09/12] famfs: Register secondary daxdevs by path (FAMFSIOC_DAXDEV_OPEN) Message-ID: References: <0100019fc572ca94-ec363dd7-3a77-484b-b4b7-f2503a0931a6-000000@email.amazonses.com> <20260803022940.75910-1-john@jagalactic.com> <0100019fc574b43d-effbc86f-dd35-4154-8c2e-17699ef0edb6-000000@email.amazonses.com> <20260806052904.GI3560084@frogsfrogsfrogs> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260806052904.GI3560084@frogsfrogsfrogs> X-Rspamd-Server: rspamout05 X-Rspamd-Queue-Id: 5165220024 X-Stat-Signature: wsdnwprc31ufep4sxch9cfq8onkesyai X-Session-Marker: 6A6F686E4067726F7665732E6E6574 X-Session-ID: U2FsdGVkX19e4VzaVm7w/scC5Jjieds+bkP0iPQMh3A= X-HE-Tag: 1786054967-655130 X-HE-Meta: U2FsdGVkX1+Qg2TeFUY/rbdu+YjAmAAK0KDZIcb6vC4OemkFW8hbbeA6wN/YOtJ/nDCy7YlzV3tkY5Wo2i6IU2lKVATqIAkRasRrEyLjebYFbND6yFIE9w9bOsmF2JQv7//QV/7D4uq3f9hx9/U1BR8ukt1NQyolYyrKlfd5RnbNC3+S9Zpnu8SH+iPcY1p/r5d9TJ0Nz4KEcsIdaxYlCA/dY7TBTWebA5m7dq7M2A4ZaqcU9GrGsruOWf1qCLRFGgX/4PpsP2gl2wZW34MXTLNcuwXjeXh+bgsjYQL0QmrY5kgeGZni99ekWDnK0prsHB0SonDbOD/cgMcaFgD61OvCSywzMVaCCpxh/5gxbzxjriJfYrSkz1JbZB0GnzdP On 26/08/05 10:29PM, Darrick J. Wong wrote: > On Mon, Aug 03, 2026 at 02:29:47AM +0000, John Groves wrote: > > From: John Groves > > > > Famfs file maps (fmaps) may reference multiple daxdevs. Before passing > > an fmap that references a new daxdev, the daxdev is pushed into the > > kernel via FAMFSIOC_DAXDEV_OPEN). This adds daxdevs to daxdev_table for > > index-based resolution from famfs extents to daxdevs. > > > > Signed-off-by: John Groves > > --- > > fs/famfs/famfs_file.c | 74 ++++++++++++++++++++++++++++++++ > > include/uapi/linux/famfs_ioctl.h | 24 +++++++++++ > > 2 files changed, 98 insertions(+) > > > > diff --git a/fs/famfs/famfs_file.c b/fs/famfs/famfs_file.c > > index e7f271ce6d03..e11a55ecf8d7 100644 > > --- a/fs/famfs/famfs_file.c > > +++ b/fs/famfs/famfs_file.c > > @@ -287,6 +287,76 @@ famfs_file_init_dax(struct file *file, void __user *arg) > > return rc; > > } > > > > +/** > > + * famfs_daxdev_open() - FAMFSIOC_DAXDEV_OPEN ioctl handler > > + * @file: any file in the famfs mount (the table is per-superblock) > > + * @arg: ptr to struct famfs_ioc_daxdev in user space > > + * > > + * Register a devdax device (identified by path) into the mount's daxdev table > > + * at the caller-specified index, so files whose extents reference that index > > + * can be mapped. The path is resolved by lookup_daxdev() - the same helper the > > + * mount uses for the primary daxdev - so every slot is resolved identically. > > + * Registering exposes raw device memory, so it requires CAP_SYS_RAWIO. > > + */ > > +static int > > +famfs_daxdev_open(struct file *file, void __user *arg) > > +{ > > + struct super_block *sb = file_inode(file)->i_sb; > > + struct famfs_fs_info *fsi = sb->s_fs_info; > > + struct famfs_ioc_daxdev dd; > > + dev_t devno; > > + char *path; > > + int rc; > > + > > + if (!capable(CAP_SYS_RAWIO)) > > + return -EPERM; > > + > > + if (copy_from_user(&dd, arg, sizeof(dd))) > > + return -EFAULT; > > + > > + /* @flags is reserved; reject non-zero so it stays available */ > > + if (dd.flags) > > + return -EINVAL; > > + > > + /* > > + * If this daxdev index is already populated there is nothing to do. > > + * The index is cluster-invariant, so a valid slot already names this > > + * device; skip the path resolution entirely. install_daxdev() rechecks > > + * ->valid under the write lock, so this is purely an optimization. > > + */ > > + scoped_guard(rwsem_read, &fsi->devlist_sem) { > > + if (dd.daxdev_index >= fsi->dax_devlist->nslots) > > + return -EINVAL; > > + if (fsi->dax_devlist->devlist[dd.daxdev_index].valid) > > + return 0; > > + } > > + > > + if (dd.daxdev_path_len == 0 || dd.daxdev_path_len >= PATH_MAX) > > + return -EINVAL; > > + > > + /* +1 so the terminating NUL is included within the bound */ > > + path = strndup_user((const char __user *)(uintptr_t)dd.daxdev_path, > > + dd.daxdev_path_len + 1); > > + if (IS_ERR(path)) > > + return PTR_ERR(path); > > + > > + rc = lookup_daxdev(path, &devno); > > + if (rc) > > + goto out; > > + > > + /* > > + * The daxdev table is allocated at mount time (for the slot-0 primary), > > + * so it is always present here; no need to allocate it. > > + */ > > + rc = famfs_install_daxdev(fsi, sb, dd.daxdev_index, devno, path); > > + if (rc) > > + pr_debug("%s: failed to install daxdev index %llu (%s)\n", > > + __func__, dd.daxdev_index, path); > > +out: > > + kfree(path); > > + return rc; > > +} > > + > > /** > > * famfs_file_ioctl() - Top-level famfs file ioctl handler > > * @file: the file > > @@ -308,6 +378,10 @@ famfs_file_ioctl(struct file *file, unsigned int cmd, unsigned long arg) > > rc = 0; > > break; > > > > + case FAMFSIOC_DAXDEV_OPEN: > > + rc = famfs_daxdev_open(file, (void __user *)arg); > > + break; > > + > > case FAMFSIOC_MAP_CREATE: > > rc = famfs_file_init_dax(file, (void __user *)arg); > > break; > > diff --git a/include/uapi/linux/famfs_ioctl.h b/include/uapi/linux/famfs_ioctl.h > > index b4eb373c1ade..751d8b033c2e 100644 > > --- a/include/uapi/linux/famfs_ioctl.h > > +++ b/include/uapi/linux/famfs_ioctl.h > > @@ -77,6 +77,29 @@ struct famfs_ioc_fmap_header { > > __u64 reserved1; > > }; > > > > +/** > > + * struct famfs_ioc_daxdev - register an additional backing daxdev by path > > + * @daxdev_index: the (cluster-invariant) index this daxdev occupies in > > + * extent dev_index fields. Index 0 is the mount-time primary. > > + * @daxdev_path: userspace pointer to the devdax device path (e.g. > > + * "/dev/dax0.0"); resolved in the kernel the same way the > > + * mount primary is. > > + * @daxdev_path_len: length of the path string, not counting the NUL. > > + * @flags: reserved; must be zero. > > + * > > + * Standalone famfs registers every daxdev by path: the mount primary comes in > > + * as the mount device name, and slots 1..n come in here. (This deliberately > > + * differs from fuse's fd-based FUSE_DEV_IOC_DAXDEV_OPEN; each side is uniform > > + * within itself.) Passing the path by pointer keeps the struct fixed-size, so > > + * longer paths never require an ABI change. > > + */ > > +struct famfs_ioc_daxdev { > > + __u64 daxdev_index; > > + __u64 daxdev_path; > > + __u32 daxdev_path_len; > > Can't we just pass an open fd to a dax device and have it use that > rather than doing its own path operations? > > --D Could do, but then there would be two ways to "open" a daxdev, since the primary is passed by path via the mount command or call. If anybody thinks that's superior enough to do it 2 different ways, let me know. Otherwise I'll leave it. Thanks! John