From: Ashish Mhetre <amhetre@nvidia.com>
To: <catalin.marinas@arm.com>, <will@kernel.org>, <corbet@lwn.net>,
<skhan@linuxfoundation.org>, <robin.murphy@arm.com>,
<joro@8bytes.org>, <nicolinc@nvidia.com>, <jgg@ziepe.ca>
Cc: <linux-arm-kernel@lists.infradead.org>,
<linux-doc@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
<iommu@lists.linux.dev>, <linux-tegra@vger.kernel.org>,
Ashish Mhetre <amhetre@nvidia.com>
Subject: [PATCH v8 0/3] iommu/arm-smmu-v3: Tegra264 invalidation workaround
Date: Wed, 22 Jul 2026 10:20:30 +0000 [thread overview]
Message-ID: <20260722102033.1122277-1-amhetre@nvidia.com> (raw)
Nvidia Tegra264 SMMUs are affected by an erratum where a TLB entry can
survive an invalidation that races with concurrent traffic targeting
the same entry. The hardware-recommended software workaround is to
issue every CFGI/TLBI command (each followed by CMD_SYNC) twice.
The second issue must execute only after the first issue's CMD_SYNC
has completed, giving the sequence:
TLBI/CFGI ... CMD_SYNC TLBI/CFGI ... CMD_SYNC
ATC_INV is not affected and must not be doubled.
The erratum is not flagged by any SMMUv3 IDR/IIDR register, so it
cannot be detected from hardware ID. Tegra264 is device-tree-only
(no ACPI/IORT support), so detection is purely by compatible string.
This series is structured as a small refactor + infrastructure + enable
sequence so that each step is reviewable in isolation:
1/3 Pure refactor (no functional change): lift the existing
force-sync conditions out of arm_smmu_cmdq_batch_add_cmd_p()
into a new arm_smmu_cmdq_batch_force_sync() helper, so that
adding another condition (in patch 2) is a one-line addition.
Authored by Nicolin Chen.
2/3 Add the workaround infrastructure without enabling it. Defines
the file-local arm_smmu_erratum_repeat_tlbi_cfgi_key static key
with an inline erratum description, the shared
arm_smmu_erratum_cmd_needs_repeating() predicate, the
arm_smmu_cmdq_issue_cmdlist() wrapper that re-issues matching
cmdlists for host-internal callers, and the batch-helper
force-sync condition. The iommufd user-invalidation path keeps
issuing user commands exactly once via the raw
__arm_smmu_cmdq_issue_cmdlist().
3/3 Enable the workaround for the existing "nvidia,tegra264-smmu"
compatible, document the erratum in silicon-errata.rst, and
report it to user space via a new
IOMMU_HW_INFO_ARM_SMMUV3_ERRATA_REPEAT_TLBI_CFGI hw_info flag so
a VMM/guest can apply the workaround to its own invalidations.
The series applies cleanly on linux-next/master (base-commit below).
Changes since v7:
- Rework the nesting/iommufd handling to report the erratum to user
space and stop repeating user-issued invalidations in the host.
- 2/3: drop arm_vsmmu_can_batch_cmd() and the iommufd batching split
it added. Expose __arm_smmu_cmdq_issue_cmdlist() and make
arm_vsmmu_cache_invalidate() call it directly.
- 3/3: add the uapi flag + enum, report it from arm_smmu_hw_info()
via a new arm_smmu_erratum_repeat_tlbi_cfgi() accessor.
- Carry Jason Gunthorpe's Reviewed-by on 1/3 and 2/3, and Nicolin
Chen's Reviewed-by on 2/3. Drop earlier tags from 3/3 since it
gained the uapi/hw_info change; re-review appreciated.
Changes since v6:
- Add #include <linux/jump_label.h> now that the static key is
defined in arm-smmu-v3.c.
- Drop the unused smmu parameter from arm_vsmmu_can_batch_cmd().
- Expand the arm_smmu_cmdq_batch_force_sync() comment to note that
batches never mix CFGI/TLBI with other commands, so checking
cmds[0] alone is enough.
- Note in 3/3 that a guest kernel enabling CMDQV on Tegra264 must
also apply this workaround, since guest-level VCMDQs issue
commands directly to the hardware.
- Carry Reviewed-by: Nicolin Chen on 2/3 and 3/3.
Changes since v5:
- Move arm_smmu_erratum_cmd_needs_repeating() into arm-smmu-v3.c
and leave a declaration-only stub in arm-smmu-v3.h. Make
arm_smmu_erratum_repeat_tlbi_cfgi_key file-local static.
- Add an inline erratum/workaround description at the static key,
referenced from arm_smmu_cmdq_batch_force_sync().
- Fix (rather than drop) the misleading !n comment above
arm_smmu_cmdq_issue_cmdlist(); keep the defensive !n guard.
- Remove the unused smmu parameter from the predicate.
- Tweak 2/3 commit-message wording ("commit" vs "patch").
Changes since v4:
- Drop ARM_SMMU_OPT_REPEAT_TLBI_CFGI entirely: the option bit was
set and read on the exact same "nvidia,tegra264-smmu" compatible
as the static key, so it added no per-instance signal that the
static key did not already carry. The predicate now gates purely
on arm_smmu_erratum_repeat_tlbi_cfgi_key.
- Reorder the series so the compatible-string detection lands
last, once all the infrastructure exists:
1/3 factor out force_sync helper (unchanged)
2/3 add static key + WAR functions (no functional change)
3/3 enable the key on nvidia,tegra264-smmu + silicon-errata
Split the old v4 "Detect" and "Issue twice" patches accordingly.
- Update the /* See ARM_SMMU_OPT_REPEAT_TLBI_CFGI */ comment inside
arm_smmu_cmdq_batch_force_sync() to reference the static key
description instead.
Changes since v3:
- Drop the cmds->num == 0 early-return so the refactor is
truly "no functional change".
- Rename ARM_SMMU_OPT_TLBI_TWICE -> ARM_SMMU_OPT_REPEAT_TLBI_CFGI
and rephrase its kdoc to be hardware-agnostic.
- Rename arm_smmu_cmd_needs_tlbi_twice() ->
arm_smmu_erratum_cmd_needs_repeating() and drop the kdoc
above it.
- Replace the explicit opcode switch with a single range check
opcode >= CMDQ_OP_CFGI_STE && opcode < CMDQ_OP_ATC_INV.
- Introduce arm_smmu_erratum_repeat_tlbi_cfgi_key static key:
the predicate gates on it first so unaffected kernels pay
only a single static_branch_unlikely() check.
- Drop the verbose Tegra264-specific comments above
arm_vsmmu_can_batch_cmd() and inside the batch helper.
- Document the erratum in
Documentation/arch/arm64/silicon-errata.rst.
- Guard the repeat path in arm_smmu_cmdq_issue_cmdlist() with
an n > 0 check so cmds[0] is never inspected on an empty
cmdlist.
- Drop the carried Reviewed-by tags now that the patch
shape has changed; re-review appreciated.
Changes since v2:
- Split into a 3-patch series (refactor / detect / apply) to keep
each step small and bisectable.
- Move the classifier to arm-smmu-v3.h as static inline so the
iommufd file can share it.
- Add arm_vsmmu_can_batch_cmd() to split iommufd batches at
"needs repeating" transitions so the per-batch decision based
on the first command stays correct under mixed user input.
- Spell out in the commit message why detection is via DT and
not via IIDR/ACPI.
Changes since v1:
- Detect the erratum from the existing "nvidia,tegra264-smmu"
compatible instead of adding a new property.
- Centralise the doubling at the CMDQ submission layer and only
apply it to CFGI/TLBI (not ATC_INV).
- Drop the binding/dtsi patches accordingly.
Ashish Mhetre (2):
iommu/arm-smmu-v3: Introduce CFGI/TLBI-repeat workaround
infrastructure
iommu/arm-smmu-v3: Enable CFGI/TLBI-repeat workaround on Tegra264
Nicolin Chen (1):
iommu/arm-smmu-v3: Factor out CMDQ batch force-sync conditions
Documentation/arch/arm64/silicon-errata.rst | 2 +
.../arm/arm-smmu-v3/arm-smmu-v3-iommufd.c | 7 +-
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 89 +++++++---
drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h | 6 ++
include/uapi/linux/iommufd.h | 12 ++-
5 files changed, 102 insertions(+), 14 deletions(-)
base-commit: 0718283ab28bc3907e10b61a6b4be6fefa1cbb2f
--
2.50.1
next reply other threads:[~2026-07-22 10:21 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-22 10:20 Ashish Mhetre [this message]
2026-07-22 10:20 ` [PATCH v8 1/3] iommu/arm-smmu-v3: Factor out CMDQ batch force-sync conditions Ashish Mhetre
2026-07-22 10:20 ` [PATCH v8 2/3] iommu/arm-smmu-v3: Introduce CFGI/TLBI-repeat workaround infrastructure Ashish Mhetre
2026-07-22 19:02 ` Nicolin Chen
2026-07-22 10:20 ` [PATCH v8 3/3] iommu/arm-smmu-v3: Enable CFGI/TLBI-repeat workaround on Tegra264 Ashish Mhetre
2026-07-22 19:16 ` Nicolin Chen
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260722102033.1122277-1-amhetre@nvidia.com \
--to=amhetre@nvidia.com \
--cc=catalin.marinas@arm.com \
--cc=corbet@lwn.net \
--cc=iommu@lists.linux.dev \
--cc=jgg@ziepe.ca \
--cc=joro@8bytes.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-tegra@vger.kernel.org \
--cc=nicolinc@nvidia.com \
--cc=robin.murphy@arm.com \
--cc=skhan@linuxfoundation.org \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.