From: Mostafa Saleh <smostafa@google.com>
To: Jason Gunthorpe <jgg@nvidia.com>
Cc: iommu@lists.linux.dev, "Joerg Roedel (AMD)" <joro@8bytes.org>,
Jean-Philippe Brucker <jpb@kernel.org>,
linux-arm-kernel@lists.infradead.org,
Robin Murphy <robin.murphy@arm.com>,
Will Deacon <will@kernel.org>,
David Matlack <dmatlack@google.com>,
Pasha Tatashin <pasha.tatashin@soleen.com>,
patches@lists.linux.dev, Pranjal Shrivastava <praan@google.com>,
Samiullah Khawaja <skhawaja@google.com>
Subject: Re: [PATCH v2 3/8] iommu/arm-smmu-v3: Optimize range invalidation for latency
Date: Thu, 13 Aug 2026 15:12:03 +0000 [thread overview]
Message-ID: <an3ew9xbnLe1dr5g@google.com> (raw)
In-Reply-To: <20260813141320.GF730363@nvidia.com>
On Thu, Aug 13, 2026 at 11:13:20AM -0300, Jason Gunthorpe wrote:
>
> > Sorry I lost track of this thread and I just saw v4.
> >
> > In the mobile space, I haven't seen an SMMUv3 that does not support
> > RIL.
>
> Oh? That's very surprising, AFAIK none of our embedded chips support it
> yet.. Even the server chips are only just getting it. Are you sure?
Yes, SMMUv3 is getting more and more common in moblie HW (opposed to
custom SoC IOMMUs). Devices that I have seen in the market in the
last couple of years have RIL. For example Pixel-10 which is currently
getting upstreamed. Also, I have a mini desktop with QCOM X1 which
have RIL.
The only SMMUv3 I have seen without RIL, is an old morello board I
have.
>
> I gather it wasn't even available in ARM IP until recently ish?
>
> > However, I have seen workloads that are really sensitive to translation
> > latency (display, camera...). And I'd be concerned about those
> > regressing.
>
> Yes, those exist, but again, they are already facing these problems if
> running without RIL.
Yes, but my point is that those typically support RIL and that change
regresses them.
>
> And, for the common case of putting something into a carve out region
> it is not so likely even an expanded RIL will intersect with a
> reserved IOVA that has a high alignment.
Not necessarily, those devices can run with a small IOVA space to
reduce the page table walk length making IOVAs quite close.
>
> At least the things we have built are calibrated to handle a TLB
> reload occasionally. The isochronous TLB's are not even sized to be
> never-miss for all cases because things like 4k media require such a
> large amount of IOVA the area cost is too high.
>
> While others can do something else you are reaching into a pretty
> narrow condition to hit a problem:
> - HW that must have a never-miss TLB to work
I am not saying that, but we shouldn't over invalidate TLBs either
when it is easy to avoid that.
> - HW that doesn't have a carve out, or has a badly aligned carve out
I do not think a carveout will help. But it's a very strong constraint
to enforce carveout on all devices specially media which are quite
complex and composite by nature.
> - A SMMU that has RIL (non RIL is already worse)
This regression only impacts RIL, otherwise it does not matter.
> - A non-isochronos workload that regularly exceeds the RIL/single
> expansion thresholds
> - Unlucky IOVA allocation that places isochronous near other
> workloads in the IOVA space.
>
It is not just luck, it depends on the IOVA space and access patterns
of the device.
> > There is a clear trade-off here as you mentioned with TLBI latency,
> > would it be make sense to make that behviour configurable from a
> > module param?
>
> I think it makes sense for a driver to indicate to the core code that
> it needs isochronous and we can do more global things like change how
> single works as well. Having an isochronous flag on the domain, for
> example, would be a good overall direction.
But according to what? It makes sense to optimize server chips,
but that should not cause over-invalidation regressions on other
hardware.
>
> I'm inclined to leave this as is and let someone come with a specific
> problematic HW, rather that try to badly guess without much
> information if it might popssibly be a problem.
>
It's not really a guess, I mentioned some examples above, that I
have seen problems of translation latencies on them.
And why not the other way around:
- Which uses cases can't handle few RIL commands?
- Why those drivers does not unmap memory with a granule/IOVA fitting
to RIL?
- Why those systems does not use FQ domains in the first place?
Thanks,
Mostafa
> Then we will know the HW and can mark the driver as I suggest above.
>
> Jason
next prev parent reply other threads:[~2026-08-13 15:12 UTC|newest]
Thread overview: 58+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-06 16:26 [PATCH v2 0/8] Organize the SMMUv3 invalidation flow so iommupt can use it Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 1/8] iommu/arm-smmu-v3: Pass the parameters for the invalidation in a struct Jason Gunthorpe
2026-07-07 3:04 ` Nicolin Chen
2026-07-07 11:18 ` Mostafa Saleh
2026-07-06 16:26 ` [PATCH v2 2/8] iommu/arm-smmu-v3: Move pgsize out of arm_smmu_inv Jason Gunthorpe
2026-07-07 3:57 ` Nicolin Chen
2026-07-07 16:15 ` Jason Gunthorpe
2026-07-07 17:21 ` Nicolin Chen
2026-07-08 18:43 ` Jason Gunthorpe
2026-07-07 11:24 ` Mostafa Saleh
2026-07-07 18:08 ` Jason Gunthorpe
2026-07-11 17:38 ` Daniel Mentz
2026-07-14 18:41 ` Jason Gunthorpe
2026-07-10 4:06 ` Daniel Mentz
2026-07-10 14:28 ` Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 3/8] iommu/arm-smmu-v3: Optimize range invalidation for latency Jason Gunthorpe
2026-07-07 7:27 ` Nicolin Chen
2026-07-07 19:13 ` Jason Gunthorpe
2026-07-07 21:07 ` Nicolin Chen
2026-07-07 11:45 ` Mostafa Saleh
2026-07-08 0:10 ` Jason Gunthorpe
2026-08-13 13:53 ` Mostafa Saleh
2026-08-13 14:13 ` Jason Gunthorpe
2026-08-13 15:12 ` Mostafa Saleh [this message]
2026-08-13 17:05 ` Jason Gunthorpe
2026-08-13 17:25 ` Robin Murphy
2026-08-13 18:01 ` Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 4/8] iommu/arm-smmu-v3: Keep track in the arm_smmu_invs if RIL is used Jason Gunthorpe
2026-07-07 7:27 ` Nicolin Chen
2026-07-07 11:46 ` Mostafa Saleh
2026-07-06 16:26 ` [PATCH v2 5/8] iommu/arm-smmu-v3: Precompute the invalidation commands Jason Gunthorpe
2026-07-07 11:52 ` Mostafa Saleh
2026-07-07 14:58 ` Jason Gunthorpe
2026-07-08 9:00 ` Mostafa Saleh
2026-07-08 13:15 ` Jason Gunthorpe
2026-07-07 20:31 ` Nicolin Chen
2026-07-09 12:07 ` Jason Gunthorpe
2026-07-09 19:10 ` Nicolin Chen
2026-07-06 16:26 ` [PATCH v2 6/8] iommu/arm-smmu-v3: Populate the tlbi at the top of the call chain Jason Gunthorpe
2026-07-07 11:57 ` Mostafa Saleh
2026-07-08 18:09 ` Jason Gunthorpe
2026-07-07 21:51 ` Nicolin Chen
2026-07-08 18:40 ` Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 7/8] iommu/arm-smmu-v3: Change how the tlbi describes the invalidation Jason Gunthorpe
2026-07-06 18:00 ` Robin Murphy
2026-07-06 19:45 ` Jason Gunthorpe
2026-07-08 1:41 ` Nicolin Chen
2026-07-08 18:27 ` Jason Gunthorpe
2026-07-08 5:29 ` Nicolin Chen
2026-07-09 18:25 ` Jason Gunthorpe
2026-07-09 21:32 ` Nicolin Chen
2026-07-06 16:26 ` [PATCH v2 8/8] iommu/arm-smmu-v3: Support the DS expansion of RIL's SCALE Jason Gunthorpe
2026-07-07 23:20 ` Nicolin Chen
2026-07-08 0:02 ` Jason Gunthorpe
2026-07-08 2:10 ` Nicolin Chen
2026-07-08 13:05 ` Jason Gunthorpe
2026-07-07 12:25 ` [PATCH v2 0/8] Organize the SMMUv3 invalidation flow so iommupt can use it Mostafa Saleh
2026-07-07 15:00 ` Jason Gunthorpe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=an3ew9xbnLe1dr5g@google.com \
--to=smostafa@google.com \
--cc=dmatlack@google.com \
--cc=iommu@lists.linux.dev \
--cc=jgg@nvidia.com \
--cc=joro@8bytes.org \
--cc=jpb@kernel.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=pasha.tatashin@soleen.com \
--cc=patches@lists.linux.dev \
--cc=praan@google.com \
--cc=robin.murphy@arm.com \
--cc=skhawaja@google.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.