From: Aneesh Kumar K.V <aneesh.kumar@kernel.org>
To: Jason Gunthorpe <jgg@nvidia.com>
Cc: "Tian, Kevin" <kevin.tian@intel.com>,
"linux-coco@lists.linux.dev" <linux-coco@lists.linux.dev>,
"iommu@lists.linux.dev" <iommu@lists.linux.dev>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
"kvm@vger.kernel.org" <kvm@vger.kernel.org>,
Alexey Kardashevskiy <aik@amd.com>,
Bjorn Helgaas <helgaas@kernel.org>,
Joerg Roedel <joro@8bytes.org>,
Jonathan Cameron <jic23@kernel.org>,
Nicolin Chen <nicolinc@nvidia.com>,
Samuel Ortiz <sameo@rivosinc.com>,
Steven Price <steven.price@arm.com>,
Suzuki K Poulose <Suzuki.Poulose@arm.com>,
Will Deacon <will@kernel.org>,
Xu Yilun <yilun.xu@linux.intel.com>,
Shameer Kolothum <shameerali.kolothum.thodi@huawei.com>,
Paolo Bonzini <pbonzini@redhat.com>
Subject: Re: [RFC PATCH v6 08/11] iommufd: Add vIOMMU provider support
Date: Tue, 29 Sep 2026 21:28:53 +0530 [thread overview]
Message-ID: <yq5ald8koxqa.fsf@kernel.org> (raw)
In-Reply-To: <20260929130643.GL1616761@nvidia.com>
Jason Gunthorpe <jgg@nvidia.com> writes:
> On Tue, Sep 29, 2026 at 06:15:39PM +0530, Aneesh Kumar K.V wrote:
>> Jason Gunthorpe <jgg@nvidia.com> writes:
>>
>> > On Tue, Sep 29, 2026 at 11:44:28AM +0530, Aneesh Kumar K.V wrote:
>> >
>> >> tsm_viommu_get_ops() runs with pci_tsm_rwsem held for read and takes a
>> >> temporary reference on the backend module (the CCA module). This keeps
>> >> the selected ops callable until vIOMMU initialization completes. iommufd
>> >> then drops the reference. This does not pin a particular TSM
>> >> registration or prevent tsm_unregister().
>> >
>> > That seems over complicated. Maybe we can't get to the sane locking I
>> > suggested earlier where TSM module is stable while a driver is bound,
>> > but we absolutely must have sane locking where we can "pin" the tsm
>> > for a pdev and it cannot be unregistered for long periods of time,
>> > such as while a viommu/vdev exists.
>> >
>> > That period should start right before getting the ops and continue to
>> > until the viommu is destroyed.
>> >
>> > No hot unplug of tsm modules while things are active.
>> >
>>
>>
>> That is essentially how it works. I decided to take the module reference
>> here and the tsm_dev reference in viommu_init() to keep the rest of
>> viommu_alloc() cleaner. That is, we have:
>>
>> struct module *owner = NULL;
>>
>> ops = tsm_viommu_get_ops(idev->dev, cmd->type, &owner);
>>
>> if (!ops) {
>> ops = iommu_dev->ops->get_viommu_ops(idev->dev, cmd->type);
>>
>> rc = ops->viommu_init(viommu, idev->dev,
>> if (rc)
>> goto out_put_hwpt;
>>
>> out_put_idev:
>> module_put(owner);
>
> if we have a get_ops we need a put_ops()..
>
>> The tsm_dev and module details are needed only by tsm_viommu, not by a
>> generic SMMU driver. For viommu_init() to take ownership of the resources
>> acquired by get_ops(), I would either need to add a viommu_info argument
>> to viommu_init(), affecting all IOMMU driver implementations, or make the
>> error handling conditional and awkward.
>
> Just don't, iommufd can hold the tsm ops if it knows it created the
> viommu through tsm.
>
>
>> >> CCA vIOMMU initialization takes a tsm_dev reference, keeping the TSM
>> >> object and its PCI/TSM resources alive until the vIOMMU is
>> >> destroyed.
>> >
>> > iommufd should do this, so long as the viommu object exists the tsm
>> > for it exists. It should not be inside tsm drivers.
>> >
>>
>> That would expose more TSM details to iommufd. The reference must also
>> be acquired under pci_tsm_rwsem.
>
> That seems like overkill.
>
>> /* Pin the current TSM and revalidate the selected vIOMMU operations. */
>> struct tsm_dev * pci_tsm_viommu_get_tsm_dev(struct pci_dev *pdev,
>> enum iommu_viommu_type type,
>> const struct iommufd_viommu_ops *expected_ops)
>> {
>> const struct pci_tsm_ops *ops;
>> const struct iommufd_viommu_ops *viommu_ops;
>> struct tsm_dev *tsm_dev;
>> int ret = -ENODEV;
>>
>> {
>> guard(rwsem_read)(&pci_tsm_rwsem);
>> if (!pdev->tsm)
>> return ERR_PTR(-ENODEV);
>> if (pdev->tsm->tsm_dev->unregistering)
>> return ERR_PTR(-ENODEV);
>> ops = to_pci_tsm_ops(pdev->tsm);
>> if (!ops->viommu_get_ops)
>> return ERR_PTR(-ENODEV);
>> tsm_dev = pdev->tsm->tsm_dev;
>> /* Unlike get_device(), this also keeps PCI/TSM resources active. */
>> if (!tsm_try_get(tsm_dev))
>> return ERR_PTR(-ENODEV);
>
> This is way too complicated for what should be a very simple scheme :(
>
> rcu_read_lock()
> tsm = rcu_derference(pdev->tsm);
> if (!tsm)
> return NULL;
> /* Prevent the module from unloading which must be the only way to
> trigger unregister */
> if (!try_module_get(tsm->ops->module))
> return NULL;
> rcu_read_unlock()
>
> And this should be exposed to iommufd..
>
> viommu_create:
> tsm = tsm_get_device(pdev)
> tsm_get_viommu_ops(tsm,...)
> [..]
>
> viommu_destroy:
> tsm_put_device(tsm)
>
> No lock, no registering FSM, just RCU free the pdev->tsm memory and do
> that module unload rcu synchronize during module unload. No hand of of
> the lifecylce to other layers.
>
> Hold the module get in iommufd inside the iommufd viommu object.
>
> It is simple and easy to understand.
>
Switching pdev->tsm to an RCU-protected pointer requires broader changes
to the existing TSM code. To move this series forward, I will continue
protecting it with pci_tsm_rwsem in the next revision, as that requires
fewer changes. We can revisit an RCU conversion later if needed?
With this approach, viommu_alloc() will do:
struct tsm_dev *tsm_dev = NULL;
tsm_dev = tsm_get_device(idev->dev);
ops = tsm_dev ? tsm_viommu_get_ops(tsm_dev, idev->dev, cmd->type) : NULL;
if (!ops) {
tsm_dev = NULL;
ops = iommu_dev->ops->get_viommu_ops(idev->dev, cmd->type);
}
viommu = (struct iommufd_viommu *)_iommufd_object_alloc_ucmd(
viommu->tsm_dev = tsm_dev;
tsm_dev = NULL;
rc = ops->viommu_init(viommu, idev->dev,....)
....
if (tsm_dev)
tsm_put_device(tsm_dev);
void iommufd_viommu_destroy(struct iommufd_object *obj)
{
..
if (viommu->tsm_dev)
tsm_put_device(viommu->tsm_dev);
...
}
struct tsm_dev *pci_tsm_get_device(struct pci_dev *pdev)
{
const struct pci_tsm_ops *ops;
struct tsm_dev *tsm_dev;
guard(rwsem_read)(&pci_tsm_rwsem);
if (!pdev->tsm)
return NULL;
tsm_dev = pdev->tsm->tsm_dev;
ops = tsm_dev->pci_ops;
if (!try_module_get(ops->owner)) // arm-cca-host
return ERR_PTR(-ENODEV);
if (!tsm_try_get(tsm_dev)) {
module_put(ops->owner);
return ERR_PTR(-ENODEV);
}
return tsm_dev;
}
>
>> viommu_ops = ops->viommu_get_ops(&pdev->dev, type);
>> if (viommu_ops == expected_ops)
>> return tsm_dev;
>> if (IS_ERR(viommu_ops))
>> ret = PTR_ERR(viommu_ops);
>> }
>
>
> This should just be try_module_get and a touch of RCU.
>
>> > ideally the module refcount handles this and the only way to trigger a
>> > tsm remove is through module unload, with no sysfs path?
>>
>> I agree. The existing code takes extra care to allow tsm_unregister().
>> However, if unloading arm-cca-host.ko is the only way to trigger
>> unregistration, as it currently is, we can avoid this complexity.
>
> I've been badly traumitized by hot unplug races, bugs and deadlock,
> let's not introduce anything like this here, there is no need.
>
Are you suggesting setting suppress_bind_attrs = true for the
arm-cca-host driver?
The driver model otherwise allows the driver to be unbound. For example,
what happens when arm-rmi device is unbound from the arm-cca-host
driver? We need to support tsm_unregister() in that case. I am also not
sure whether there are other paths that can call tsm_unregister().
Currently, we register the cleanup callback via
tsm_dev = tsm_register(&sdev->dev, &cca_link_pci_ops);
ret = devm_add_action_or_reset(&sdev->dev, cca_link_tsm_remove, tsm_dev);
if (ret)
-aneesh
next prev parent reply other threads:[~2026-09-29 15:59 UTC|newest]
Thread overview: 72+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 14:01 [RFC PATCH v6 00/11] iommufd: Infrastructure for vIOMMU creation for confidential guests and guest TSM requests Aneesh Kumar K.V (Arm)
2026-09-17 14:01 ` [RFC PATCH v6 01/11] vfio: cache KVM VM file references instead of raw struct kvm pointers Aneesh Kumar K.V (Arm)
2026-09-24 19:41 ` Jason Gunthorpe
2026-09-30 7:19 ` Aneesh Kumar K.V
2026-09-17 14:01 ` [RFC PATCH v6 02/11] vfio: cdev: Reject duplicate bind before updating KVM file Aneesh Kumar K.V (Arm)
2026-09-24 7:53 ` Tian, Kevin
2026-09-25 5:49 ` Aneesh Kumar K.V
2026-09-17 14:01 ` [RFC PATCH v6 03/11] iommufd/device: Associate KVM file pointer with iommufd_device Aneesh Kumar K.V (Arm)
2026-09-17 14:01 ` [RFC PATCH v6 04/11] iommufd/viommu: Keep a reference to the KVM file Aneesh Kumar K.V (Arm)
2026-09-30 13:28 ` Vasant Hegde
2026-10-02 5:26 ` Aneesh Kumar K.V
2026-09-17 14:01 ` [RFC PATCH v6 05/11] iommu: Add a helper to validate a vIOMMU parent Aneesh Kumar K.V (Arm)
2026-09-24 19:41 ` Jason Gunthorpe
2026-09-25 5:48 ` Aneesh Kumar K.V
2026-09-25 12:23 ` Jason Gunthorpe
2026-09-28 10:36 ` Aneesh Kumar K.V
2026-09-28 12:11 ` Jason Gunthorpe
2026-09-28 15:39 ` Aneesh Kumar K.V
2026-09-28 16:17 ` Jason Gunthorpe
2026-09-28 18:08 ` Jacob Pan
2026-09-28 18:20 ` Jason Gunthorpe
2026-09-28 22:24 ` Jacob Pan
2026-09-28 23:03 ` Jason Gunthorpe
2026-09-29 5:55 ` Jacob Pan
2026-09-29 12:30 ` Jason Gunthorpe
2026-09-29 23:15 ` Jacob Pan
2026-09-29 23:30 ` Jason Gunthorpe
2026-09-17 14:01 ` [RFC PATCH v6 06/11] iommu: Add a helper to query vIOMMU hardware parameters Aneesh Kumar K.V (Arm)
2026-09-24 19:41 ` Jason Gunthorpe
2026-09-25 5:59 ` Aneesh Kumar K.V
2026-09-25 12:29 ` Jason Gunthorpe
2026-09-17 14:01 ` [RFC PATCH v6 07/11] coco: tsm: Expose active-user lifetime references Aneesh Kumar K.V (Arm)
2026-09-17 14:01 ` [RFC PATCH v6 08/11] iommufd: Add vIOMMU provider support Aneesh Kumar K.V (Arm)
2026-09-24 19:41 ` Jason Gunthorpe
2026-09-25 6:08 ` Aneesh Kumar K.V
2026-09-25 12:39 ` Jason Gunthorpe
2026-09-28 3:41 ` Tian, Kevin
2026-09-29 6:14 ` Aneesh Kumar K.V
2026-09-29 12:17 ` Jason Gunthorpe
2026-09-29 12:45 ` Aneesh Kumar K.V
2026-09-29 13:06 ` Jason Gunthorpe
2026-09-29 15:58 ` Aneesh Kumar K.V [this message]
2026-09-29 19:10 ` Jason Gunthorpe
2026-09-28 10:51 ` Aneesh Kumar K.V
2026-09-17 14:01 ` [RFC PATCH v6 09/11] iommufd: Add the vdevice TSM request ioctl Aneesh Kumar K.V (Arm)
2026-09-18 13:09 ` Alexey Kardashevskiy
2026-09-18 13:13 ` Jason Gunthorpe
2026-09-24 19:41 ` Jason Gunthorpe
2026-09-30 8:03 ` Aneesh Kumar K.V
2026-09-30 13:26 ` Vasant Hegde
2026-09-30 14:03 ` Jason Gunthorpe
2026-09-30 14:08 ` Jason Gunthorpe
2026-09-17 14:01 ` [RFC PATCH v6 10/11] PCI/TSM: Remove the legacy guest request interface Aneesh Kumar K.V (Arm)
2026-09-17 14:01 ` [RFC PATCH v6 11/11] PCI/TSM: Add reference-counted contexts for vdevice providers Aneesh Kumar K.V (Arm)
2026-09-24 8:17 ` Tian, Kevin
2026-09-24 19:41 ` Jason Gunthorpe
2026-09-25 8:15 ` Aneesh Kumar K.V
2026-09-28 18:47 ` Sonang Patel
2026-09-28 23:08 ` Jason Gunthorpe
2026-10-02 6:14 ` Aneesh Kumar K.V
2026-10-02 13:06 ` Jason Gunthorpe
2026-09-17 14:17 ` [RFC PATCH v6 00/11] iommufd: Infrastructure for vIOMMU creation for confidential guests and guest TSM requests Aneesh Kumar K.V
2026-09-24 7:48 ` Tian, Kevin
2026-09-24 19:20 ` Jason Gunthorpe
2026-09-28 3:35 ` Tian, Kevin
2026-09-28 13:08 ` Jason Gunthorpe
2026-09-25 8:29 ` Aneesh Kumar K.V
2026-09-28 3:41 ` Tian, Kevin
2026-09-28 3:55 ` Tian, Kevin
2026-09-24 19:09 ` Jason Gunthorpe
2026-09-25 6:46 ` Aneesh Kumar K.V
2026-09-25 12:45 ` Jason Gunthorpe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=yq5ald8koxqa.fsf@kernel.org \
--to=aneesh.kumar@kernel.org \
--cc=Suzuki.Poulose@arm.com \
--cc=aik@amd.com \
--cc=helgaas@kernel.org \
--cc=iommu@lists.linux.dev \
--cc=jgg@nvidia.com \
--cc=jic23@kernel.org \
--cc=joro@8bytes.org \
--cc=kevin.tian@intel.com \
--cc=kvm@vger.kernel.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-kernel@vger.kernel.org \
--cc=nicolinc@nvidia.com \
--cc=pbonzini@redhat.com \
--cc=sameo@rivosinc.com \
--cc=shameerali.kolothum.thodi@huawei.com \
--cc=steven.price@arm.com \
--cc=will@kernel.org \
--cc=yilun.xu@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.