From: Jason Gunthorpe <jgg@nvidia.com>
To: Robin Murphy <robin.murphy@arm.com>
Cc: Mostafa Saleh <smostafa@google.com>,
iommu@lists.linux.dev, "Joerg Roedel (AMD)" <joro@8bytes.org>,
Jean-Philippe Brucker <jpb@kernel.org>,
linux-arm-kernel@lists.infradead.org,
Will Deacon <will@kernel.org>,
David Matlack <dmatlack@google.com>,
Pasha Tatashin <pasha.tatashin@soleen.com>,
patches@lists.linux.dev, Pranjal Shrivastava <praan@google.com>,
Samiullah Khawaja <skhawaja@google.com>
Subject: Re: [PATCH v2 3/8] iommu/arm-smmu-v3: Optimize range invalidation for latency
Date: Thu, 13 Aug 2026 15:01:02 -0300 [thread overview]
Message-ID: <20260813180102.GH730363@nvidia.com> (raw)
In-Reply-To: <1791d7da-ecd9-4475-8f27-c8857477b442@arm.com>
On Thu, Aug 13, 2026 at 06:25:59PM +0100, Robin Murphy wrote:
> > > Yes, those exist, but again, they are already facing these problems if
> > > running without RIL.
> >
> > Yes, but my point is that those typically support RIL and that change
> > regresses them.
>
> Not even that - the main concern is over-invalidation of adjacent in-use
> buffers due to rounding up; non-RIL absolutely does not have that issue and
> never has.
It certainly does! SMMUv3 got an invalidate all path a while back
because doing single for >> MB's of IOVA effectively soft lockups the
system - especially with SVA.
We set the cut off at ~2M which matches when the CPU goes to
invalidate all. That's 512 commands max of single TBLIs before we just
dump the entire TLB. That's a huge over invalidation.
> Note that with RIL, even precise invalidation should only actually need at
> most two commands (excepting absurd off-the-scale sizes) - the trick is to
> consider that they can overlap.
Oh that's really interesting, I never thought about doing it like
that. It is way better than the algorithm that is there right now.
Let me try it, it seems like it would make everyone happy.
> Conversely though, if we really did have a demonstrable need to minimise the
> number of commands issued then we should probably also not bother with the
> TTL hint nor splitting leaf ranges from non-leaf, such that we can gather
> pretty much any unmap into a single command.
That is what I am doing. The series converting to iommupt also changes
to use iommupt style gathers which default to combining everything
into one gather.
One gather maps to one tlbi in this series and it turns into one
RIL. The hints/etc are used if the gather happens to be compatible,
otherwise the single RIL is still pushed un-hinted.
> This is probably something we'll end up wanting some kind of tunable
> behaviour for, given that SMMUv3 hardware is going to be spanning an
> increasingly wide range of use-cases with increasingly opposing requirements
> - it's certainly more than just "hypervisor or not". For now, though, I'm
> also not buying a dubious latency argument based on apparently no real-world
> data other than "I think"...
We have data already showing that large numbers of invalidation
commands cause soft lockups, and we had to fix SMMU for this.
This is tied into the SVA path remember, latency directly effects mm
application benchmarks and we have been consistently working toward
bounding and reducing SVA invalidation latency. We actually did some
studies recently and invalidation latency had a major impact on real
application metrics. We did not build HW vCMDQ support to reduce
invalidation latences in VMs for no reason. It is not "I think".
Thanks,
Jason
next prev parent reply other threads:[~2026-08-13 18:01 UTC|newest]
Thread overview: 58+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-06 16:26 [PATCH v2 0/8] Organize the SMMUv3 invalidation flow so iommupt can use it Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 1/8] iommu/arm-smmu-v3: Pass the parameters for the invalidation in a struct Jason Gunthorpe
2026-07-07 3:04 ` Nicolin Chen
2026-07-07 11:18 ` Mostafa Saleh
2026-07-06 16:26 ` [PATCH v2 2/8] iommu/arm-smmu-v3: Move pgsize out of arm_smmu_inv Jason Gunthorpe
2026-07-07 3:57 ` Nicolin Chen
2026-07-07 16:15 ` Jason Gunthorpe
2026-07-07 17:21 ` Nicolin Chen
2026-07-08 18:43 ` Jason Gunthorpe
2026-07-07 11:24 ` Mostafa Saleh
2026-07-07 18:08 ` Jason Gunthorpe
2026-07-11 17:38 ` Daniel Mentz
2026-07-14 18:41 ` Jason Gunthorpe
2026-07-10 4:06 ` Daniel Mentz
2026-07-10 14:28 ` Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 3/8] iommu/arm-smmu-v3: Optimize range invalidation for latency Jason Gunthorpe
2026-07-07 7:27 ` Nicolin Chen
2026-07-07 19:13 ` Jason Gunthorpe
2026-07-07 21:07 ` Nicolin Chen
2026-07-07 11:45 ` Mostafa Saleh
2026-07-08 0:10 ` Jason Gunthorpe
2026-08-13 13:53 ` Mostafa Saleh
2026-08-13 14:13 ` Jason Gunthorpe
2026-08-13 15:12 ` Mostafa Saleh
2026-08-13 17:05 ` Jason Gunthorpe
2026-08-13 17:25 ` Robin Murphy
2026-08-13 18:01 ` Jason Gunthorpe [this message]
2026-07-06 16:26 ` [PATCH v2 4/8] iommu/arm-smmu-v3: Keep track in the arm_smmu_invs if RIL is used Jason Gunthorpe
2026-07-07 7:27 ` Nicolin Chen
2026-07-07 11:46 ` Mostafa Saleh
2026-07-06 16:26 ` [PATCH v2 5/8] iommu/arm-smmu-v3: Precompute the invalidation commands Jason Gunthorpe
2026-07-07 11:52 ` Mostafa Saleh
2026-07-07 14:58 ` Jason Gunthorpe
2026-07-08 9:00 ` Mostafa Saleh
2026-07-08 13:15 ` Jason Gunthorpe
2026-07-07 20:31 ` Nicolin Chen
2026-07-09 12:07 ` Jason Gunthorpe
2026-07-09 19:10 ` Nicolin Chen
2026-07-06 16:26 ` [PATCH v2 6/8] iommu/arm-smmu-v3: Populate the tlbi at the top of the call chain Jason Gunthorpe
2026-07-07 11:57 ` Mostafa Saleh
2026-07-08 18:09 ` Jason Gunthorpe
2026-07-07 21:51 ` Nicolin Chen
2026-07-08 18:40 ` Jason Gunthorpe
2026-07-06 16:26 ` [PATCH v2 7/8] iommu/arm-smmu-v3: Change how the tlbi describes the invalidation Jason Gunthorpe
2026-07-06 18:00 ` Robin Murphy
2026-07-06 19:45 ` Jason Gunthorpe
2026-07-08 1:41 ` Nicolin Chen
2026-07-08 18:27 ` Jason Gunthorpe
2026-07-08 5:29 ` Nicolin Chen
2026-07-09 18:25 ` Jason Gunthorpe
2026-07-09 21:32 ` Nicolin Chen
2026-07-06 16:26 ` [PATCH v2 8/8] iommu/arm-smmu-v3: Support the DS expansion of RIL's SCALE Jason Gunthorpe
2026-07-07 23:20 ` Nicolin Chen
2026-07-08 0:02 ` Jason Gunthorpe
2026-07-08 2:10 ` Nicolin Chen
2026-07-08 13:05 ` Jason Gunthorpe
2026-07-07 12:25 ` [PATCH v2 0/8] Organize the SMMUv3 invalidation flow so iommupt can use it Mostafa Saleh
2026-07-07 15:00 ` Jason Gunthorpe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260813180102.GH730363@nvidia.com \
--to=jgg@nvidia.com \
--cc=dmatlack@google.com \
--cc=iommu@lists.linux.dev \
--cc=joro@8bytes.org \
--cc=jpb@kernel.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=pasha.tatashin@soleen.com \
--cc=patches@lists.linux.dev \
--cc=praan@google.com \
--cc=robin.murphy@arm.com \
--cc=skhawaja@google.com \
--cc=smostafa@google.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.