From: Gerd Bayer <gbayer@linux.ibm.com>
To: Matthew Rosato <mjrosato@linux.ibm.com>,
sashiko-reviews@lists.linux.dev,
Niklas Schnelle <schnelle@linux.ibm.com>
Cc: Alexander Gordeev <agordeev@linux.ibm.com>,
Vasily Gorbik <gor@linux.ibm.com>,
linux-s390@vger.kernel.org,
Christian Borntraeger <borntraeger@linux.ibm.com>,
Heiko Carstens <hca@linux.ibm.com>,
Gerd Bayer <gbayer@linux.ibm.com>
Subject: Re: [PATCH 2/3] s390/pci: Use 32-bit fh outside of inline asm
Date: Thu, 08 Oct 2026 10:50:26 +0200 [thread overview]
Message-ID: <f082e43f024a3aa4a897b388b346eff4edea37ec.camel@linux.ibm.com> (raw)
In-Reply-To: <9e3894fa-0a1e-432b-ba1d-76ea34d4f648@linux.ibm.com>
On Wed, 2026-10-07 at 12:21 -0400, Matthew Rosato wrote:
> On 10/7/26 10:46 AM, Gerd Bayer wrote:
> > On Wed, 2026-10-07 at 11:19 +0000, sashiko-bot@kernel.org wrote:
> > > Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> > > - [High] s390/pci: infinite loop in zpci_refresh_trans on CC=2
> > > --
> > >
> > > commit 1c8706bb963f0d551da7d8535f804ca9f9b4305a
> > > Author: Gerd Bayer <gbayer@linux.ibm.com>
> > >
> > > s390/pci: Use 32-bit fh outside of inline asm
> > >
> > > zpci_refresh_trans() declared the PCI function handle as a u64 type just
> > > because the actual RPCIT machine instruction required that in a 64bit
> > > register, while the function handle is architected as a 32bit type and
> > > defined as such in the zpci_dev struct.
> > >
> > > > diff --git a/arch/s390/pci/pci_insn.c b/arch/s390/pci/pci_insn.c
> > > > index 2ba787e2304b5..053afb3e7a990 100644
> > > > --- a/arch/s390/pci/pci_insn.c
> > > > +++ b/arch/s390/pci/pci_insn.c
> > > > @@ -117,13 +117,14 @@ static inline u8 __rpcit(u64 fn, u64 addr, u64 range, u8 *status)
> > > > return CC_TRANSFORM(cc);
> > > > }
> > > >
> > > > -int zpci_refresh_trans(u64 fn, u64 addr, u64 range)
> > > > +int zpci_refresh_trans(u32 fh, u64 addr, u64 range)
> > > > {
> > > > + u64 req = ZPCI_CREATE_REQ(fh, 0, 0);
> > > > bool retried = false;
> > > > u8 cc, status;
> > > >
> > > > do {
> > > > - cc = __rpcit(fn, addr, range, &status);
> > > > + cc = __rpcit(req, addr, range, &status);
> > > > if (cc == 2) {
> > > > udelay(ZPCI_INSN_BUSY_DELAY);
> > >
> > > [Severity: High]
> > > Does this retry loop risk causing an infinite loop or CPU stall when handling
> > > large DMA spaces?
> > >
>
> This should really have been flagged as a pre-existing issue vs a new
> regression, it is not a behavior introduced (or changed) by your patch.
Yes, I agree. This patch didn't touch/remove the capability to pass
back a modified addr or range.
>
> > > When the s390x RPCIT instruction sets CC=2, it indicates that it cannot complete
> > > the operation in the allotted time. It then updates the address and range
> > > registers to reflect the remaining uncompleted portion.
> >
> > I've re-read the relevant ISA documentation a couple times now. I don't
> > find any mention of address and range being modified after RPCIT was
> > executed - regardless of condition code or status.
> >
> > Maybe @Niklas or @Matt have an idea where this allegation of "CC=2
> > means partial completion and one should retry with an updated
> > iova/range" comes from?
>
> FWIW, it's not coming from QEMU -- we don't even have a CC=2 return case
> for RPCIT today, and we never update the iova/range on return.
>
> Perhaps, lacking access to the ISA documentation, it made an assumption
> that the range would be updated to shorten the next RPCIT as this would
> allow us to converge towards a smaller and smaller range each attempt to
> potentially not hit the busy condition eventually - sounds logical anyway?
>
> But I also find nothing that indicates the range will/should be updated
> on CC=2 (or otherwise) from RPCIT; so it does just sound like a bad
> assumption.
Thanks for the confirmation.
> >
> > > However, in __rpcit(), addr_range is passed using an input-only constraint ("d"
> > > instead of "+d"), and the arguments are passed by value:
> > >
> > > arch/s390/pci/pci_insn.c:__rpcit() {
> > > union register_pair addr_range = {.even = addr, .odd = range};
> > > ...
> > > : [addr_range] "d" (addr_range.pair)
> > > }
> > >
> > > Because the hardware's partial progress updates to addr and range are discarded,
> > > this loop continually re-invokes __rpcit() with the unmodified original
> > > addr and range.
>
> That matches your observation: we indeed aren't taking updates to the
> range in __rpcit(). But as you say: there aren't supposed to be any.
>
> > >
> > > If scanning the initial portion consistently exhausts the hardware time slice
> > > (such as during IOMMU TLB flushes on devices with large DMA spaces in
> > > zpci_refresh_all()), will the operation fail to advance and loop infinitely?
> >
>
> That concern sounds valid at face value; if we tried something and it
> couldn't be completed in time, from a linux perspective we are simply
> trying it again without changing anything and hoping everything works
> for the best this time.
>
> > I see that all the PCI related instructions in this file may retry
> > indefinitely (with delay) on CC=2. However, the architecture guarantees
> > that at some point the instruction will end with any of the other
> > condition codes.
>
> I think that is the rub. Sashiko couldn't possibly know that such an
> architecture guarantee exists.
>
> I suspect the question is also a bit theoretical: what happens if the
> range is so big that it's impossible for the RPCIT to process it before
> the CC2 trigger point due simply to how big the range provided is? Then
> you're guaranteed to hit it every time unless you 'chip away' at it by
> shrinking the range after each attempt.
>
> I would assume/hope that the SDMA-EDMA range limitation we impose on the
> aperture size already ensures that it is easily possible to handle a
> RPCIT from SDMA thru EDMA without tripping CC=2.
> So then the CC=2 case becomes purely situational vs a guaranteed result
> even for the largest possible RPCIT request. That already makes it far
> more reasonable to simply try it again.
>
> So: assuming that architecture guarantee stands (we know RPCIT can
> possibly complete for the largest possible range SDMA-EDMA, and we have
> some architected guarantee that the loop will be broken otherwise e.g.
> architecture says it won't keep giving us CC2s forever) then it sounds
> like nothing to fix, but maybe worth a comment in a future patch that
> explains these guarantees, based on what you find in the ISA.
>
> I suppose we can consider whether linux should have its own redundancy
> here or not in addition to those guarantees. That sounds well beyond
> the scope of this series and goes back to what I mentioned at the start:
> this is pre-existing bevahior and should not hold up this patch either way.
I went back to commit cd24834130ac ("s390/pci: base support") from 2012
when RPCIT was introduced: The potentially endless loop on CC=2 was
introduced at the very beginning.
I found explicit words in the ISA specifying that "no other action is
taken" when the instruction ends with CC=2. IIRC, CC=2 was introduced
to signal a program that there are internal inhibitors present that
prohibit to even start the requested "refresh" - or other PCI
instructions. I still have to search for a statement that guarantees at
the architecture level that CC=2 won't be presented forever. When I
find that, I'll think about how to frame that into a future patch.
>
> Thanks,
> Matt
Thank you,
Gerd
next prev parent reply other threads:[~2026-10-08 8:50 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-07 11:07 [PATCH 0/3] s390/pci: Updates to PCI insn tracing Gerd Bayer
2026-10-07 11:07 ` [PATCH 1/3] s390/pci: Adjust alignment in insn trace Gerd Bayer
2026-10-07 11:14 ` sashiko-bot
2026-10-07 13:10 ` Gerd Bayer
2026-10-07 14:36 ` Niklas Schnelle
2026-10-07 11:07 ` [PATCH 2/3] s390/pci: Use 32-bit fh outside of inline asm Gerd Bayer
2026-10-07 11:19 ` sashiko-bot
2026-10-07 14:46 ` Gerd Bayer
2026-10-07 16:21 ` Matthew Rosato
2026-10-08 8:50 ` Gerd Bayer [this message]
2026-10-08 9:28 ` Niklas Schnelle
2026-10-07 11:07 ` [PATCH 3/3] s390/pci: Add function handle to RPCIT insn trace Gerd Bayer
2026-10-07 11:17 ` sashiko-bot
2026-10-07 13:12 ` Gerd Bayer
2026-10-07 14:38 ` Niklas Schnelle
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=f082e43f024a3aa4a897b388b346eff4edea37ec.camel@linux.ibm.com \
--to=gbayer@linux.ibm.com \
--cc=agordeev@linux.ibm.com \
--cc=borntraeger@linux.ibm.com \
--cc=gor@linux.ibm.com \
--cc=hca@linux.ibm.com \
--cc=linux-s390@vger.kernel.org \
--cc=mjrosato@linux.ibm.com \
--cc=sashiko-reviews@lists.linux.dev \
--cc=schnelle@linux.ibm.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox