From: Jason Gunthorpe <jgg@nvidia.com>
To: Pranjal Shrivastava <praan@google.com>
Cc: Catalin Marinas <catalin.marinas@arm.com>,
Jonathan Corbet <corbet@lwn.net>,
iommu@lists.linux.dev, "Joerg Roedel (AMD)" <joro@8bytes.org>,
Jean-Philippe Brucker <jpb@kernel.org>,
linux-arm-kernel@lists.infradead.org, linux-doc@vger.kernel.org,
Mark Rutland <mark.rutland@arm.com>,
Randy Dunlap <rdunlap@infradead.org>,
Robin Murphy <robin.murphy@arm.com>,
Shuah Khan <skhan@linuxfoundation.org>,
Will Deacon <will@kernel.org>,
David Matlack <dmatlack@google.com>,
Jean-Philippe Brucker <jean-philippe@linaro.org>,
Jonathan Cameron <Jonathan.Cameron@huawei.com>,
Nicolin Chen <nicolinc@nvidia.com>,
Pasha Tatashin <pasha.tatashin@soleen.com>,
patches@lists.linux.dev, Samiullah Khawaja <skhawaja@google.com>,
Mostafa Saleh <smostafa@google.com>,
stable@vger.kernel.org,
Vijayanand Jitta <vijayanand.jitta@oss.qualcomm.com>
Subject: Re: [PATCH v7 4/9] iommu/arm-smmu-v3: Optimize range invalidation for latency
Date: Mon, 28 Sep 2026 10:50:07 -0300 [thread overview]
Message-ID: <20260928135007.GA1616761@nvidia.com> (raw)
In-Reply-To: <arpohmaJTPEEI7mi@google.com>
On Mon, Sep 28, 2026 at 01:15:50PM +0000, Pranjal Shrivastava wrote:
> > -/* Used by non INV_TYPE_ATS* invalidations */
> > -static void arm_smmu_inv_to_cmdq_batch(struct arm_smmu_inv *inv,
> > +/*
> > + * One TLBI command per IOTLB entry, assuming the entries are all at least
> > + * iopte_granule sized. Returns false if too many commands would be needed which
> > + * indicates too high a latency. The threshold is similar to MAX_DVM_OPS in
> > + * arch/arm64/include/asm/tlbflush.h for the 4k PAGE_SIZE.
> > + */
> > +static bool arm_smmu_cmdq_batch_add_single(struct arm_smmu_device *smmu,
> > + struct arm_smmu_cmdq_batch *cmds,
> > + struct arm_smmu_cmd *cmd,
> > + struct arm_smmu_tlbi *tlbi)
> > +{
> > + unsigned long num_ops = tlbi->size / tlbi->iopte_size;
> > + unsigned long iova = tlbi->iova;
> > + unsigned long i;
> > +
> > + if (!num_ops || num_ops > 512)
> > + return false;
> > +
>
> Should this be ">= 512"? The old arm_smmu_inv_size_too_big() and
> arm64's __flush_tlb_range_limit_excess() both fall back to a full
> invalidation at exactly 512:
>
> return pages >= (MAX_DVM_OPS * stride) >> PAGE_SHIFT;
>
> 512 is what a walk flush could generate. For example, when
> __arm_lpae_unmap() frees a last-level table on a 4K granule,
> it calls io_pgtable_tlb_flush_walk(iova, SZ_2M, SZ_4K), hence,
> num_ops == 512.
>
> Before this patch that was a single TLBI_NH_ASID/TLBI_S12_VMALL, now
> it becomes 512 VA TLBIs + CMD_SYNC for every 2M table freed on a
> non-RIL SMMU. The same check carries into arm_smmu_tlbi_calc_single()
> later in the series.
I had set it deliberately like that so that a table flush would issue
singles. That is largely based around the feedback from before that we
shouldn't over invalidate.
We've been insensitive to that issue in the past, and I'm going to
argue the >= of today's code is a bug as it effectively made alot of
iopgtable actions turn into new full invalidations, while originally
they were range bound singles.
So I view this choice as a return to the historical behavior before we
started to fix the soft lockup issues. I guess it crept it during one
of the refactorings where we merged the SVA limitation into the main
flow.
Granted I didn't notice this detail that the current code was 511 not
512, I will add a remark to the commit message:
The end result is any gather is converted into either:
- One invalidate all
- One or two range invalidation operations
- At most 512 single invalidation ops
The current code switches at 511, but this turns every table removal
into a full invalidation. Change to 512 to try to minimize the amount
of full invalidation.
Thanks,
Jason
next prev parent reply other threads:[~2026-09-28 13:50 UTC|newest]
Thread overview: 23+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-21 23:55 [PATCH v7 0/9] Organize the SMMUv3 invalidation flow so iommupt can use it Jason Gunthorpe
2026-09-21 23:55 ` [PATCH v7 1/9] iommu/arm-smmu-v3: Handle ARM erratum for CONT under invalidation with SVA Jason Gunthorpe
2026-09-28 11:35 ` Pranjal Shrivastava
2026-09-21 23:55 ` [PATCH v7 2/9] iommu/arm-smmu-v3: Pass the parameters for the invalidation in a struct Jason Gunthorpe
2026-09-28 11:34 ` Pranjal Shrivastava
2026-09-21 23:55 ` [PATCH v7 3/9] iommu/arm-smmu-v3: Move pgsize out of arm_smmu_inv Jason Gunthorpe
2026-09-28 11:36 ` Pranjal Shrivastava
2026-09-21 23:55 ` [PATCH v7 4/9] iommu/arm-smmu-v3: Optimize range invalidation for latency Jason Gunthorpe
2026-09-28 13:15 ` Pranjal Shrivastava
2026-09-28 13:50 ` Jason Gunthorpe [this message]
2026-09-28 16:37 ` Pranjal Shrivastava
2026-09-21 23:55 ` [PATCH v7 5/9] iommu/arm-smmu-v3: Keep track in arm_smmu_invs if range invalidation is used Jason Gunthorpe
2026-09-28 14:56 ` Pranjal Shrivastava
2026-09-21 23:55 ` [PATCH v7 6/9] iommu/arm-smmu-v3: Precompute the invalidation commands Jason Gunthorpe
2026-09-28 16:39 ` Pranjal Shrivastava
2026-09-21 23:55 ` [PATCH v7 7/9] iommu/arm-smmu-v3: Populate the tlbi at the top of the call chain Jason Gunthorpe
2026-09-28 17:05 ` Pranjal Shrivastava
2026-09-28 18:11 ` Jason Gunthorpe
2026-09-21 23:55 ` [PATCH v7 8/9] iommu/arm-smmu-v3: Change how the tlbi describes the invalidation Jason Gunthorpe
2026-09-28 18:25 ` Pranjal Shrivastava
2026-09-21 23:55 ` [PATCH v7 9/9] iommu/arm-smmu-v3: Support the DS expansion of range invalidation SCALE Jason Gunthorpe
2026-09-28 18:44 ` Pranjal Shrivastava
2026-09-28 19:45 ` [PATCH v7 0/9] Organize the SMMUv3 invalidation flow so iommupt can use it Pranjal Shrivastava
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928135007.GA1616761@nvidia.com \
--to=jgg@nvidia.com \
--cc=Jonathan.Cameron@huawei.com \
--cc=catalin.marinas@arm.com \
--cc=corbet@lwn.net \
--cc=dmatlack@google.com \
--cc=iommu@lists.linux.dev \
--cc=jean-philippe@linaro.org \
--cc=joro@8bytes.org \
--cc=jpb@kernel.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-doc@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=nicolinc@nvidia.com \
--cc=pasha.tatashin@soleen.com \
--cc=patches@lists.linux.dev \
--cc=praan@google.com \
--cc=rdunlap@infradead.org \
--cc=robin.murphy@arm.com \
--cc=skhan@linuxfoundation.org \
--cc=skhawaja@google.com \
--cc=smostafa@google.com \
--cc=stable@vger.kernel.org \
--cc=vijayanand.jitta@oss.qualcomm.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox