From: Aneesh Kumar K.V <aneesh.kumar@kernel.org>
To: Jason Gunthorpe <jgg@nvidia.com>
Cc: "Tian, Kevin" <kevin.tian@intel.com>,
"linux-coco@lists.linux.dev" <linux-coco@lists.linux.dev>,
"iommu@lists.linux.dev" <iommu@lists.linux.dev>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
"kvm@vger.kernel.org" <kvm@vger.kernel.org>,
Alexey Kardashevskiy <aik@amd.com>,
Bjorn Helgaas <helgaas@kernel.org>,
Joerg Roedel <joro@8bytes.org>,
Jonathan Cameron <jic23@kernel.org>,
Nicolin Chen <nicolinc@nvidia.com>,
Samuel Ortiz <sameo@rivosinc.com>,
Steven Price <steven.price@arm.com>,
Suzuki K Poulose <Suzuki.Poulose@arm.com>,
Will Deacon <will@kernel.org>,
Xu Yilun <yilun.xu@linux.intel.com>,
Shameer Kolothum <shameerali.kolothum.thodi@huawei.com>,
Paolo Bonzini <pbonzini@redhat.com>
Subject: Re: [RFC PATCH v6 08/11] iommufd: Add vIOMMU provider support
Date: Tue, 29 Sep 2026 18:15:39 +0530 [thread overview]
Message-ID: <yq5aqzicp6oc.fsf@kernel.org> (raw)
In-Reply-To: <20260929121723.GI1616761@nvidia.com>
Jason Gunthorpe <jgg@nvidia.com> writes:
> On Tue, Sep 29, 2026 at 11:44:28AM +0530, Aneesh Kumar K.V wrote:
>
>> tsm_viommu_get_ops() runs with pci_tsm_rwsem held for read and takes a
>> temporary reference on the backend module (the CCA module). This keeps
>> the selected ops callable until vIOMMU initialization completes. iommufd
>> then drops the reference. This does not pin a particular TSM
>> registration or prevent tsm_unregister().
>
> That seems over complicated. Maybe we can't get to the sane locking I
> suggested earlier where TSM module is stable while a driver is bound,
> but we absolutely must have sane locking where we can "pin" the tsm
> for a pdev and it cannot be unregistered for long periods of time,
> such as while a viommu/vdev exists.
>
> That period should start right before getting the ops and continue to
> until the viommu is destroyed.
>
> No hot unplug of tsm modules while things are active.
>
That is essentially how it works. I decided to take the module reference
here and the tsm_dev reference in viommu_init() to keep the rest of
viommu_alloc() cleaner. That is, we have:
struct module *owner = NULL;
ops = tsm_viommu_get_ops(idev->dev, cmd->type, &owner);
if (!ops) {
ops = iommu_dev->ops->get_viommu_ops(idev->dev, cmd->type);
rc = ops->viommu_init(viommu, idev->dev,
if (rc)
goto out_put_hwpt;
out_put_idev:
module_put(owner);
The tsm_dev and module details are needed only by tsm_viommu, not by a
generic SMMU driver. For viommu_init() to take ownership of the resources
acquired by get_ops(), I would either need to add a viommu_info argument
to viommu_init(), affecting all IOMMU driver implementations, or make the
error handling conditional and awkward.
What I have now is that get_ops() pins the module, so the CCA module
cannot disappear after get_ops() returns. We retain that reference until
viommu_init() completes, and viommu_init() establishes the long-term
pins internally. This keeps the API simpler.
>
>> CCA vIOMMU initialization takes a tsm_dev reference, keeping the TSM
>> object and its PCI/TSM resources alive until the vIOMMU is
>> destroyed.
>
> iommufd should do this, so long as the viommu object exists the tsm
> for it exists. It should not be inside tsm drivers.
>
That would expose more TSM details to iommufd. The reference must also
be acquired under pci_tsm_rwsem.
/* Pin the current TSM and revalidate the selected vIOMMU operations. */
struct tsm_dev * pci_tsm_viommu_get_tsm_dev(struct pci_dev *pdev,
enum iommu_viommu_type type,
const struct iommufd_viommu_ops *expected_ops)
{
const struct pci_tsm_ops *ops;
const struct iommufd_viommu_ops *viommu_ops;
struct tsm_dev *tsm_dev;
int ret = -ENODEV;
{
guard(rwsem_read)(&pci_tsm_rwsem);
if (!pdev->tsm)
return ERR_PTR(-ENODEV);
if (pdev->tsm->tsm_dev->unregistering)
return ERR_PTR(-ENODEV);
ops = to_pci_tsm_ops(pdev->tsm);
if (!ops->viommu_get_ops)
return ERR_PTR(-ENODEV);
tsm_dev = pdev->tsm->tsm_dev;
/* Unlike get_device(), this also keeps PCI/TSM resources active. */
if (!tsm_try_get(tsm_dev))
return ERR_PTR(-ENODEV);
viommu_ops = ops->viommu_get_ops(&pdev->dev, type);
if (viommu_ops == expected_ops)
return tsm_dev;
if (IS_ERR(viommu_ops))
ret = PTR_ERR(viommu_ops);
}
tsm_put(tsm_dev);
return ERR_PTR(ret);
}
EXPORT_SYMBOL_GPL(pci_tsm_viommu_get_tsm_dev);
>
>> The initialization callback also takes a reference on the CCA module so
>> that the vIOMMU callbacks remain available after iommufd drops the
>> temporary discovery reference described above.
>
> The initial pin should do this since it is the only way to prevent
> unregistration.
>
see above
>
>> For a link TSM-connected device, a vdevice holds a pci_tsm_context
>> reference. The context holds device references and increments PF0's
>> context_users under the PF0 mutex. PCI/TSM disconnect checks that count
>> under the same mutex and returns -EBUSY while contexts remain. The
>> context is released during vdevice teardown. This prevents link
>> disconnect while a vdevice is active without blocking tsm_unregister().
>
> This one seems reasonable, but not sure a mutex is needed on top of a
> a simple refcount scheme.
>
The PF0 mutex is already used by PCI/TSM to serialize some of these
operations. I can double-check whether it is necessary here.
>
>> A new unregistering state is added to tsm_dev. tsm_unregister() sets it,
>> unregisters the class device, and drops the registration reference. New
>> vIOMMU allocations, vdevice contexts, and PCI TSM connect/lock
>> operations reject the TSM once this state is set. Existing users retain
>> their references and can be torn down normally, so tsm_unregister() does
>> not need to wait for them. PCI/TSM teardown occurs when the last active
>> tsm_dev reference is dropped.
>
> This seems over complicated, tsm unregistration should be made
> impossible while it is not able to complete. We shouldn't need the
> complexity of states here when we don't need to support tsm
> hot-unplug.
>
>
> ideally the module refcount handles this and the only way to trigger a
> tsm remove is through module unload, with no sysfs path?
>
I agree. The existing code takes extra care to allow tsm_unregister().
However, if unloading arm-cca-host.ko is the only way to trigger
unregistration, as it currently is, we can avoid this complexity.
-aneesh
next prev parent reply other threads:[~2026-09-29 12:45 UTC|newest]
Thread overview: 72+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 14:01 [RFC PATCH v6 00/11] iommufd: Infrastructure for vIOMMU creation for confidential guests and guest TSM requests Aneesh Kumar K.V (Arm)
2026-09-17 14:01 ` [RFC PATCH v6 01/11] vfio: cache KVM VM file references instead of raw struct kvm pointers Aneesh Kumar K.V (Arm)
2026-09-24 19:41 ` Jason Gunthorpe
2026-09-30 7:19 ` Aneesh Kumar K.V
2026-09-17 14:01 ` [RFC PATCH v6 02/11] vfio: cdev: Reject duplicate bind before updating KVM file Aneesh Kumar K.V (Arm)
2026-09-24 7:53 ` Tian, Kevin
2026-09-25 5:49 ` Aneesh Kumar K.V
2026-09-17 14:01 ` [RFC PATCH v6 03/11] iommufd/device: Associate KVM file pointer with iommufd_device Aneesh Kumar K.V (Arm)
2026-09-17 14:01 ` [RFC PATCH v6 04/11] iommufd/viommu: Keep a reference to the KVM file Aneesh Kumar K.V (Arm)
2026-09-30 13:28 ` Vasant Hegde
2026-10-02 5:26 ` Aneesh Kumar K.V
2026-09-17 14:01 ` [RFC PATCH v6 05/11] iommu: Add a helper to validate a vIOMMU parent Aneesh Kumar K.V (Arm)
2026-09-24 19:41 ` Jason Gunthorpe
2026-09-25 5:48 ` Aneesh Kumar K.V
2026-09-25 12:23 ` Jason Gunthorpe
2026-09-28 10:36 ` Aneesh Kumar K.V
2026-09-28 12:11 ` Jason Gunthorpe
2026-09-28 15:39 ` Aneesh Kumar K.V
2026-09-28 16:17 ` Jason Gunthorpe
2026-09-28 18:08 ` Jacob Pan
2026-09-28 18:20 ` Jason Gunthorpe
2026-09-28 22:24 ` Jacob Pan
2026-09-28 23:03 ` Jason Gunthorpe
2026-09-29 5:55 ` Jacob Pan
2026-09-29 12:30 ` Jason Gunthorpe
2026-09-29 23:15 ` Jacob Pan
2026-09-29 23:30 ` Jason Gunthorpe
2026-09-17 14:01 ` [RFC PATCH v6 06/11] iommu: Add a helper to query vIOMMU hardware parameters Aneesh Kumar K.V (Arm)
2026-09-24 19:41 ` Jason Gunthorpe
2026-09-25 5:59 ` Aneesh Kumar K.V
2026-09-25 12:29 ` Jason Gunthorpe
2026-09-17 14:01 ` [RFC PATCH v6 07/11] coco: tsm: Expose active-user lifetime references Aneesh Kumar K.V (Arm)
2026-09-17 14:01 ` [RFC PATCH v6 08/11] iommufd: Add vIOMMU provider support Aneesh Kumar K.V (Arm)
2026-09-24 19:41 ` Jason Gunthorpe
2026-09-25 6:08 ` Aneesh Kumar K.V
2026-09-25 12:39 ` Jason Gunthorpe
2026-09-28 3:41 ` Tian, Kevin
2026-09-29 6:14 ` Aneesh Kumar K.V
2026-09-29 12:17 ` Jason Gunthorpe
2026-09-29 12:45 ` Aneesh Kumar K.V [this message]
2026-09-29 13:06 ` Jason Gunthorpe
2026-09-29 15:58 ` Aneesh Kumar K.V
2026-09-29 19:10 ` Jason Gunthorpe
2026-09-28 10:51 ` Aneesh Kumar K.V
2026-09-17 14:01 ` [RFC PATCH v6 09/11] iommufd: Add the vdevice TSM request ioctl Aneesh Kumar K.V (Arm)
2026-09-18 13:09 ` Alexey Kardashevskiy
2026-09-18 13:13 ` Jason Gunthorpe
2026-09-24 19:41 ` Jason Gunthorpe
2026-09-30 8:03 ` Aneesh Kumar K.V
2026-09-30 13:26 ` Vasant Hegde
2026-09-30 14:03 ` Jason Gunthorpe
2026-09-30 14:08 ` Jason Gunthorpe
2026-09-17 14:01 ` [RFC PATCH v6 10/11] PCI/TSM: Remove the legacy guest request interface Aneesh Kumar K.V (Arm)
2026-09-17 14:01 ` [RFC PATCH v6 11/11] PCI/TSM: Add reference-counted contexts for vdevice providers Aneesh Kumar K.V (Arm)
2026-09-24 8:17 ` Tian, Kevin
2026-09-24 19:41 ` Jason Gunthorpe
2026-09-25 8:15 ` Aneesh Kumar K.V
2026-09-28 18:47 ` Sonang Patel
2026-09-28 23:08 ` Jason Gunthorpe
2026-10-02 6:14 ` Aneesh Kumar K.V
2026-10-02 13:06 ` Jason Gunthorpe
2026-09-17 14:17 ` [RFC PATCH v6 00/11] iommufd: Infrastructure for vIOMMU creation for confidential guests and guest TSM requests Aneesh Kumar K.V
2026-09-24 7:48 ` Tian, Kevin
2026-09-24 19:20 ` Jason Gunthorpe
2026-09-28 3:35 ` Tian, Kevin
2026-09-28 13:08 ` Jason Gunthorpe
2026-09-25 8:29 ` Aneesh Kumar K.V
2026-09-28 3:41 ` Tian, Kevin
2026-09-28 3:55 ` Tian, Kevin
2026-09-24 19:09 ` Jason Gunthorpe
2026-09-25 6:46 ` Aneesh Kumar K.V
2026-09-25 12:45 ` Jason Gunthorpe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=yq5aqzicp6oc.fsf@kernel.org \
--to=aneesh.kumar@kernel.org \
--cc=Suzuki.Poulose@arm.com \
--cc=aik@amd.com \
--cc=helgaas@kernel.org \
--cc=iommu@lists.linux.dev \
--cc=jgg@nvidia.com \
--cc=jic23@kernel.org \
--cc=joro@8bytes.org \
--cc=kevin.tian@intel.com \
--cc=kvm@vger.kernel.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-kernel@vger.kernel.org \
--cc=nicolinc@nvidia.com \
--cc=pbonzini@redhat.com \
--cc=sameo@rivosinc.com \
--cc=shameerali.kolothum.thodi@huawei.com \
--cc=steven.price@arm.com \
--cc=will@kernel.org \
--cc=yilun.xu@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.