Linux-ARM-Kernel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Mostafa Saleh <smostafa@google.com>
To: Jason Gunthorpe <jgg@nvidia.com>
Cc: iommu@lists.linux.dev, "Joerg Roedel (AMD)" <joro@8bytes.org>,
	Jean-Philippe Brucker <jpb@kernel.org>,
	linux-arm-kernel@lists.infradead.org,
	Robin Murphy <robin.murphy@arm.com>,
	Will Deacon <will@kernel.org>,
	David Matlack <dmatlack@google.com>,
	Pasha Tatashin <pasha.tatashin@soleen.com>,
	patches@lists.linux.dev, Pranjal Shrivastava <praan@google.com>,
	Samiullah Khawaja <skhawaja@google.com>
Subject: Re: [PATCH v2 3/8] iommu/arm-smmu-v3: Optimize range invalidation for latency
Date: Thu, 13 Aug 2026 15:12:03 +0000	[thread overview]
Message-ID: <an3ew9xbnLe1dr5g@google.com> (raw)
In-Reply-To: <20260813141320.GF730363@nvidia.com>

On Thu, Aug 13, 2026 at 11:13:20AM -0300, Jason Gunthorpe wrote:
> 
> > Sorry I lost track of this thread and I just saw v4.
> > 
> > In the mobile space, I haven't seen an SMMUv3 that does not support
> > RIL.
> 
> Oh? That's very surprising, AFAIK none of our embedded chips support it
> yet.. Even the server chips are only just getting it. Are you sure?

Yes, SMMUv3 is getting more and more common in moblie HW (opposed to
custom SoC IOMMUs). Devices that I have seen in the market in the
last couple of years have RIL. For example Pixel-10 which is currently
getting upstreamed. Also, I have a mini desktop with QCOM X1 which
have RIL.

The only SMMUv3 I have seen without RIL, is an old morello board I
have.

> 
> I gather it wasn't even available in ARM IP until recently ish?
> 
> > However, I have seen workloads that are really sensitive to translation
> > latency (display, camera...). And I'd be concerned about those
> > regressing.
> 
> Yes, those exist, but again, they are already facing these problems if
> running without RIL.

Yes, but my point is that those typically support RIL and that change
regresses them.

> 
> And, for the common case of putting something into a carve out region
> it is not so likely even an expanded RIL will intersect with a
> reserved IOVA that has a high alignment.

Not necessarily, those devices can run with a small IOVA space to
reduce the page table walk length making IOVAs quite close.

> 
> At least the things we have built are calibrated to handle a TLB
> reload occasionally. The isochronous TLB's are not even sized to be
> never-miss for all cases because things like 4k media require such a
> large amount of IOVA the area cost is too high.
> 
> While others can do something else you are reaching into a pretty
> narrow condition to hit a problem:
>  - HW that must have a never-miss TLB to work

I am not saying that, but we shouldn't over invalidate TLBs either
when it is easy to avoid that.

>  - HW that doesn't have a carve out, or has a badly aligned carve out

I do not think a carveout will help. But it's a very strong constraint
to enforce carveout on all devices specially media which are quite
complex and composite by nature.

>  - A SMMU that has RIL (non RIL is already worse)

This regression only impacts RIL, otherwise it does not matter.

>  - A non-isochronos workload that regularly exceeds the RIL/single
>    expansion thresholds
>  - Unlucky IOVA allocation that places isochronous near other
>    workloads in the IOVA space.
> 

It is not just luck, it depends on the IOVA space and access patterns
of the device.

> > There is a clear trade-off here as you mentioned with TLBI latency,
> > would it be make sense to make that behviour configurable from a
> > module param?
> 
> I think it makes sense for a driver to indicate to the core code that
> it needs isochronous and we can do more global things like change how
> single works as well. Having an isochronous flag on the domain, for
> example, would be a good overall direction.

But according to what? It makes sense to optimize server chips,
but that should not cause over-invalidation regressions on other
hardware.

> 
> I'm inclined to leave this as is and let someone come with a specific
> problematic HW, rather that try to badly guess without much
> information if it might popssibly be a problem.
> 

It's not really a guess, I mentioned some examples above, that I
have seen problems of translation latencies on them.

And why not the other way around:
- Which uses cases can't handle few RIL commands?
- Why those drivers does not unmap memory with a granule/IOVA fitting
  to RIL?
- Why those systems does not use FQ domains in the first place?

Thanks,
Mostafa

> Then we will know the HW and can mark the driver as I suggest above.
> 
> Jason


  reply	other threads:[~2026-08-13 15:12 UTC|newest]

Thread overview: 55+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-06 16:26 [PATCH v2 0/8] Organize the SMMUv3 invalidation flow so iommupt can use it Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 1/8] iommu/arm-smmu-v3: Pass the parameters for the invalidation in a struct Jason Gunthorpe
2026-07-07  3:04   ` Nicolin Chen
2026-07-07 11:18   ` Mostafa Saleh
2026-07-06 16:26 ` [PATCH v2 2/8] iommu/arm-smmu-v3: Move pgsize out of arm_smmu_inv Jason Gunthorpe
2026-07-07  3:57   ` Nicolin Chen
2026-07-07 16:15     ` Jason Gunthorpe
2026-07-07 17:21       ` Nicolin Chen
2026-07-08 18:43         ` Jason Gunthorpe
2026-07-07 11:24   ` Mostafa Saleh
2026-07-07 18:08     ` Jason Gunthorpe
2026-07-11 17:38       ` Daniel Mentz
2026-07-14 18:41         ` Jason Gunthorpe
2026-07-10  4:06   ` Daniel Mentz
2026-07-10 14:28     ` Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 3/8] iommu/arm-smmu-v3: Optimize range invalidation for latency Jason Gunthorpe
2026-07-07  7:27   ` Nicolin Chen
2026-07-07 19:13     ` Jason Gunthorpe
2026-07-07 21:07       ` Nicolin Chen
2026-07-07 11:45   ` Mostafa Saleh
2026-07-08  0:10     ` Jason Gunthorpe
2026-08-13 13:53       ` Mostafa Saleh
2026-08-13 14:13         ` Jason Gunthorpe
2026-08-13 15:12           ` Mostafa Saleh [this message]
2026-07-06 16:26 ` [PATCH v2 4/8] iommu/arm-smmu-v3: Keep track in the arm_smmu_invs if RIL is used Jason Gunthorpe
2026-07-07  7:27   ` Nicolin Chen
2026-07-07 11:46   ` Mostafa Saleh
2026-07-06 16:26 ` [PATCH v2 5/8] iommu/arm-smmu-v3: Precompute the invalidation commands Jason Gunthorpe
2026-07-07 11:52   ` Mostafa Saleh
2026-07-07 14:58     ` Jason Gunthorpe
2026-07-08  9:00       ` Mostafa Saleh
2026-07-08 13:15         ` Jason Gunthorpe
2026-07-07 20:31   ` Nicolin Chen
2026-07-09 12:07     ` Jason Gunthorpe
2026-07-09 19:10       ` Nicolin Chen
2026-07-06 16:26 ` [PATCH v2 6/8] iommu/arm-smmu-v3: Populate the tlbi at the top of the call chain Jason Gunthorpe
2026-07-07 11:57   ` Mostafa Saleh
2026-07-08 18:09     ` Jason Gunthorpe
2026-07-07 21:51   ` Nicolin Chen
2026-07-08 18:40     ` Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 7/8] iommu/arm-smmu-v3: Change how the tlbi describes the invalidation Jason Gunthorpe
2026-07-06 18:00   ` Robin Murphy
2026-07-06 19:45     ` Jason Gunthorpe
2026-07-08  1:41   ` Nicolin Chen
2026-07-08 18:27     ` Jason Gunthorpe
2026-07-08  5:29   ` Nicolin Chen
2026-07-09 18:25     ` Jason Gunthorpe
2026-07-09 21:32       ` Nicolin Chen
2026-07-06 16:26 ` [PATCH v2 8/8] iommu/arm-smmu-v3: Support the DS expansion of RIL's SCALE Jason Gunthorpe
2026-07-07 23:20   ` Nicolin Chen
2026-07-08  0:02     ` Jason Gunthorpe
2026-07-08  2:10       ` Nicolin Chen
2026-07-08 13:05         ` Jason Gunthorpe
2026-07-07 12:25 ` [PATCH v2 0/8] Organize the SMMUv3 invalidation flow so iommupt can use it Mostafa Saleh
2026-07-07 15:00   ` Jason Gunthorpe

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=an3ew9xbnLe1dr5g@google.com \
    --to=smostafa@google.com \
    --cc=dmatlack@google.com \
    --cc=iommu@lists.linux.dev \
    --cc=jgg@nvidia.com \
    --cc=joro@8bytes.org \
    --cc=jpb@kernel.org \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=pasha.tatashin@soleen.com \
    --cc=patches@lists.linux.dev \
    --cc=praan@google.com \
    --cc=robin.murphy@arm.com \
    --cc=skhawaja@google.com \
    --cc=will@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox