From: Mostafa Saleh <smostafa@google.com>
To: Jason Gunthorpe <jgg@nvidia.com>
Cc: iommu@lists.linux.dev, "Joerg Roedel (AMD)" <joro@8bytes.org>,
Jean-Philippe Brucker <jpb@kernel.org>,
linux-arm-kernel@lists.infradead.org,
Robin Murphy <robin.murphy@arm.com>,
Will Deacon <will@kernel.org>,
David Matlack <dmatlack@google.com>,
Pasha Tatashin <pasha.tatashin@soleen.com>,
patches@lists.linux.dev, Pranjal Shrivastava <praan@google.com>,
Samiullah Khawaja <skhawaja@google.com>
Subject: Re: [PATCH v2 3/8] iommu/arm-smmu-v3: Optimize range invalidation for latency
Date: Thu, 13 Aug 2026 15:12:03 +0000 [thread overview]
Message-ID: <an3ew9xbnLe1dr5g@google.com> (raw)
In-Reply-To: <20260813141320.GF730363@nvidia.com>
On Thu, Aug 13, 2026 at 11:13:20AM -0300, Jason Gunthorpe wrote:
>
> > Sorry I lost track of this thread and I just saw v4.
> >
> > In the mobile space, I haven't seen an SMMUv3 that does not support
> > RIL.
>
> Oh? That's very surprising, AFAIK none of our embedded chips support it
> yet.. Even the server chips are only just getting it. Are you sure?
Yes, SMMUv3 is getting more and more common in moblie HW (opposed to
custom SoC IOMMUs). Devices that I have seen in the market in the
last couple of years have RIL. For example Pixel-10 which is currently
getting upstreamed. Also, I have a mini desktop with QCOM X1 which
have RIL.
The only SMMUv3 I have seen without RIL, is an old morello board I
have.
>
> I gather it wasn't even available in ARM IP until recently ish?
>
> > However, I have seen workloads that are really sensitive to translation
> > latency (display, camera...). And I'd be concerned about those
> > regressing.
>
> Yes, those exist, but again, they are already facing these problems if
> running without RIL.
Yes, but my point is that those typically support RIL and that change
regresses them.
>
> And, for the common case of putting something into a carve out region
> it is not so likely even an expanded RIL will intersect with a
> reserved IOVA that has a high alignment.
Not necessarily, those devices can run with a small IOVA space to
reduce the page table walk length making IOVAs quite close.
>
> At least the things we have built are calibrated to handle a TLB
> reload occasionally. The isochronous TLB's are not even sized to be
> never-miss for all cases because things like 4k media require such a
> large amount of IOVA the area cost is too high.
>
> While others can do something else you are reaching into a pretty
> narrow condition to hit a problem:
> - HW that must have a never-miss TLB to work
I am not saying that, but we shouldn't over invalidate TLBs either
when it is easy to avoid that.
> - HW that doesn't have a carve out, or has a badly aligned carve out
I do not think a carveout will help. But it's a very strong constraint
to enforce carveout on all devices specially media which are quite
complex and composite by nature.
> - A SMMU that has RIL (non RIL is already worse)
This regression only impacts RIL, otherwise it does not matter.
> - A non-isochronos workload that regularly exceeds the RIL/single
> expansion thresholds
> - Unlucky IOVA allocation that places isochronous near other
> workloads in the IOVA space.
>
It is not just luck, it depends on the IOVA space and access patterns
of the device.
> > There is a clear trade-off here as you mentioned with TLBI latency,
> > would it be make sense to make that behviour configurable from a
> > module param?
>
> I think it makes sense for a driver to indicate to the core code that
> it needs isochronous and we can do more global things like change how
> single works as well. Having an isochronous flag on the domain, for
> example, would be a good overall direction.
But according to what? It makes sense to optimize server chips,
but that should not cause over-invalidation regressions on other
hardware.
>
> I'm inclined to leave this as is and let someone come with a specific
> problematic HW, rather that try to badly guess without much
> information if it might popssibly be a problem.
>
It's not really a guess, I mentioned some examples above, that I
have seen problems of translation latencies on them.
And why not the other way around:
- Which uses cases can't handle few RIL commands?
- Why those drivers does not unmap memory with a granule/IOVA fitting
to RIL?
- Why those systems does not use FQ domains in the first place?
Thanks,
Mostafa
> Then we will know the HW and can mark the driver as I suggest above.
>
> Jason
next prev parent reply other threads:[~2026-08-13 15:12 UTC|newest]
Thread overview: 58+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-06 16:26 [PATCH v2 0/8] Organize the SMMUv3 invalidation flow so iommupt can use it Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 1/8] iommu/arm-smmu-v3: Pass the parameters for the invalidation in a struct Jason Gunthorpe
2026-07-07 3:04 ` Nicolin Chen
2026-07-07 11:18 ` Mostafa Saleh
2026-07-06 16:26 ` [PATCH v2 2/8] iommu/arm-smmu-v3: Move pgsize out of arm_smmu_inv Jason Gunthorpe
2026-07-07 3:57 ` Nicolin Chen
2026-07-07 16:15 ` Jason Gunthorpe
2026-07-07 17:21 ` Nicolin Chen
2026-07-08 18:43 ` Jason Gunthorpe
2026-07-07 11:24 ` Mostafa Saleh
2026-07-07 18:08 ` Jason Gunthorpe
2026-07-11 17:38 ` Daniel Mentz
2026-07-14 18:41 ` Jason Gunthorpe
2026-07-10 4:06 ` Daniel Mentz
2026-07-10 14:28 ` Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 3/8] iommu/arm-smmu-v3: Optimize range invalidation for latency Jason Gunthorpe
2026-07-07 7:27 ` Nicolin Chen
2026-07-07 19:13 ` Jason Gunthorpe
2026-07-07 21:07 ` Nicolin Chen
2026-07-07 11:45 ` Mostafa Saleh
2026-07-08 0:10 ` Jason Gunthorpe
2026-08-13 13:53 ` Mostafa Saleh
2026-08-13 14:13 ` Jason Gunthorpe
2026-08-13 15:12 ` Mostafa Saleh [this message]
2026-08-13 17:05 ` Jason Gunthorpe
2026-08-13 17:25 ` Robin Murphy
2026-08-13 18:01 ` Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 4/8] iommu/arm-smmu-v3: Keep track in the arm_smmu_invs if RIL is used Jason Gunthorpe
2026-07-07 7:27 ` Nicolin Chen
2026-07-07 11:46 ` Mostafa Saleh
2026-07-06 16:26 ` [PATCH v2 5/8] iommu/arm-smmu-v3: Precompute the invalidation commands Jason Gunthorpe
2026-07-07 11:52 ` Mostafa Saleh
2026-07-07 14:58 ` Jason Gunthorpe
2026-07-08 9:00 ` Mostafa Saleh
2026-07-08 13:15 ` Jason Gunthorpe
2026-07-07 20:31 ` Nicolin Chen
2026-07-09 12:07 ` Jason Gunthorpe
2026-07-09 19:10 ` Nicolin Chen
2026-07-06 16:26 ` [PATCH v2 6/8] iommu/arm-smmu-v3: Populate the tlbi at the top of the call chain Jason Gunthorpe
2026-07-07 11:57 ` Mostafa Saleh
2026-07-08 18:09 ` Jason Gunthorpe
2026-07-07 21:51 ` Nicolin Chen
2026-07-08 18:40 ` Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 7/8] iommu/arm-smmu-v3: Change how the tlbi describes the invalidation Jason Gunthorpe
2026-07-06 18:00 ` Robin Murphy
2026-07-06 19:45 ` Jason Gunthorpe
2026-07-08 1:41 ` Nicolin Chen
2026-07-08 18:27 ` Jason Gunthorpe
2026-07-08 5:29 ` Nicolin Chen
2026-07-09 18:25 ` Jason Gunthorpe
2026-07-09 21:32 ` Nicolin Chen
2026-07-06 16:26 ` [PATCH v2 8/8] iommu/arm-smmu-v3: Support the DS expansion of RIL's SCALE Jason Gunthorpe
2026-07-07 23:20 ` Nicolin Chen
2026-07-08 0:02 ` Jason Gunthorpe
2026-07-08 2:10 ` Nicolin Chen
2026-07-08 13:05 ` Jason Gunthorpe
2026-07-07 12:25 ` [PATCH v2 0/8] Organize the SMMUv3 invalidation flow so iommupt can use it Mostafa Saleh
2026-07-07 15:00 ` Jason Gunthorpe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=an3ew9xbnLe1dr5g@google.com \
--to=smostafa@google.com \
--cc=dmatlack@google.com \
--cc=iommu@lists.linux.dev \
--cc=jgg@nvidia.com \
--cc=joro@8bytes.org \
--cc=jpb@kernel.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=pasha.tatashin@soleen.com \
--cc=patches@lists.linux.dev \
--cc=praan@google.com \
--cc=robin.murphy@arm.com \
--cc=skhawaja@google.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox