Linux s390 Architecture development
 help / color / mirror / Atom feed
From: Gerd Bayer <gbayer@linux.ibm.com>
To: Matthew Rosato <mjrosato@linux.ibm.com>,
	sashiko-reviews@lists.linux.dev,
	Niklas Schnelle <schnelle@linux.ibm.com>
Cc: Alexander Gordeev <agordeev@linux.ibm.com>,
	Vasily Gorbik	 <gor@linux.ibm.com>,
	linux-s390@vger.kernel.org,
	Christian Borntraeger	 <borntraeger@linux.ibm.com>,
	Heiko Carstens <hca@linux.ibm.com>,
	Gerd Bayer <gbayer@linux.ibm.com>
Subject: Re: [PATCH 2/3] s390/pci: Use 32-bit fh outside of inline asm
Date: Thu, 08 Oct 2026 10:50:26 +0200	[thread overview]
Message-ID: <f082e43f024a3aa4a897b388b346eff4edea37ec.camel@linux.ibm.com> (raw)
In-Reply-To: <9e3894fa-0a1e-432b-ba1d-76ea34d4f648@linux.ibm.com>

On Wed, 2026-10-07 at 12:21 -0400, Matthew Rosato wrote:
> On 10/7/26 10:46 AM, Gerd Bayer wrote:
> > On Wed, 2026-10-07 at 11:19 +0000, sashiko-bot@kernel.org wrote:
> > > Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> > > - [High] s390/pci: infinite loop in zpci_refresh_trans on CC=2
> > > --
> > > 
> > > commit 1c8706bb963f0d551da7d8535f804ca9f9b4305a
> > > Author: Gerd Bayer <gbayer@linux.ibm.com>
> > > 
> > > s390/pci: Use 32-bit fh outside of inline asm
> > >     
> > > zpci_refresh_trans() declared the PCI function handle as a u64 type just
> > > because the actual RPCIT machine instruction required that in a 64bit
> > > register, while the function handle is architected as a 32bit type and
> > > defined as such in the zpci_dev struct.
> > > 
> > > > diff --git a/arch/s390/pci/pci_insn.c b/arch/s390/pci/pci_insn.c
> > > > index 2ba787e2304b5..053afb3e7a990 100644
> > > > --- a/arch/s390/pci/pci_insn.c
> > > > +++ b/arch/s390/pci/pci_insn.c
> > > > @@ -117,13 +117,14 @@ static inline u8 __rpcit(u64 fn, u64 addr, u64 range, u8 *status)
> > > >  	return CC_TRANSFORM(cc);
> > > >  }
> > > >  
> > > > -int zpci_refresh_trans(u64 fn, u64 addr, u64 range)
> > > > +int zpci_refresh_trans(u32 fh, u64 addr, u64 range)
> > > >  {
> > > > +	u64 req = ZPCI_CREATE_REQ(fh, 0, 0);
> > > >  	bool retried = false;
> > > >  	u8 cc, status;
> > > >  
> > > >  	do {
> > > > -		cc = __rpcit(fn, addr, range, &status);
> > > > +		cc = __rpcit(req, addr, range, &status);
> > > >  		if (cc == 2) {
> > > >  			udelay(ZPCI_INSN_BUSY_DELAY);
> > > 
> > > [Severity: High]
> > > Does this retry loop risk causing an infinite loop or CPU stall when handling
> > > large DMA spaces?
> > > 
> 
> This should really have been flagged as a pre-existing issue vs a new
> regression, it is not a behavior introduced (or changed) by your patch.

Yes, I agree. This patch didn't touch/remove the capability to pass
back a modified addr or range.

> 
> > > When the s390x RPCIT instruction sets CC=2, it indicates that it cannot complete
> > > the operation in the allotted time. It then updates the address and range
> > > registers to reflect the remaining uncompleted portion.
> > 
> > I've re-read the relevant ISA documentation a couple times now. I don't
> > find any mention of address and range being modified after RPCIT was
> > executed - regardless of condition code or status.
> > 
> > Maybe @Niklas or @Matt have an idea where this allegation of "CC=2
> > means partial completion and one should retry with an updated
> > iova/range" comes from?
> 
> FWIW, it's not coming from QEMU -- we don't even have a CC=2 return case
> for RPCIT today, and we never update the iova/range on return.
> 
> Perhaps, lacking access to the ISA documentation, it made an assumption
> that the range would be updated to shorten the next RPCIT as this would
> allow us to converge towards a smaller and smaller range each attempt to
> potentially not hit the busy condition eventually - sounds logical anyway?
> 
> But I also find nothing that indicates the range will/should be updated
> on CC=2 (or otherwise) from RPCIT; so it does just sound like a bad
> assumption.

Thanks for the confirmation.

> > 
> > > However, in __rpcit(), addr_range is passed using an input-only constraint ("d"
> > > instead of "+d"), and the arguments are passed by value:
> > > 
> > > arch/s390/pci/pci_insn.c:__rpcit() {
> > > 	union register_pair addr_range = {.even = addr, .odd = range};
> > > 	...
> > > 		: [addr_range] "d" (addr_range.pair)
> > > }
> > > 
> > > Because the hardware's partial progress updates to addr and range are discarded,
> > > this loop continually re-invokes __rpcit() with the unmodified original
> > > addr and range.
> 
> That matches your observation:  we indeed aren't taking updates to the
> range in __rpcit().  But as you say: there aren't supposed to be any.
> 
> > > 
> > > If scanning the initial portion consistently exhausts the hardware time slice
> > > (such as during IOMMU TLB flushes on devices with large DMA spaces in
> > > zpci_refresh_all()), will the operation fail to advance and loop infinitely?
> > 
> 
> That concern sounds valid at face value; if we tried something and it
> couldn't be completed in time, from a linux perspective we are simply
> trying it again without changing anything and hoping everything works
> for the best this time.
> 
> > I see that all the PCI related instructions in this file may retry
> > indefinitely (with delay) on CC=2. However, the architecture guarantees
> > that at some point the instruction will end with any of the other
> > condition codes.
> 
> I think that is the rub.  Sashiko couldn't possibly know that such an
> architecture guarantee exists.
> 
> I suspect the question is also a bit theoretical: what happens if the
> range is so big that it's impossible for the RPCIT to process it before
> the CC2 trigger point due simply to how big the range provided is?  Then
> you're guaranteed to hit it every time unless you 'chip away' at it by
> shrinking the range after each attempt.
> 
> I would assume/hope that the SDMA-EDMA range limitation we impose on the
> aperture size already ensures that it is easily possible to handle a
> RPCIT from SDMA thru EDMA without tripping CC=2.
> So then the CC=2 case becomes purely situational vs a guaranteed result
> even for the largest possible RPCIT request.  That already makes it far
> more reasonable to simply try it again.
> 
> So: assuming that architecture guarantee stands (we know RPCIT can
> possibly complete for the largest possible range SDMA-EDMA, and we have
> some architected guarantee that the loop will be broken otherwise e.g.
> architecture says it won't keep giving us CC2s forever) then it sounds
> like nothing to fix, but maybe worth a comment in a future patch that
> explains these guarantees, based on what you find in the ISA.
> 
> I suppose we can consider whether linux should have its own redundancy
> here or not in addition to those guarantees.  That sounds well beyond
> the scope of this series and goes back to what I mentioned at the start:
> this is pre-existing bevahior and should not hold up this patch either way.

I went back to commit cd24834130ac ("s390/pci: base support") from 2012
when RPCIT was introduced: The potentially endless loop on CC=2 was
introduced at the very beginning.

I found explicit words in the ISA specifying that "no other action is
taken" when the instruction ends with CC=2. IIRC, CC=2 was introduced
to signal a program that there are internal inhibitors present that
prohibit to even start the requested "refresh" - or other PCI
instructions. I still have to search for a statement that guarantees at
the architecture level that CC=2 won't be presented forever. When I
find that, I'll think about how to frame that into a future patch.

> 
> Thanks,
> Matt

Thank you,
Gerd

  reply	other threads:[~2026-10-08  8:50 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-07 11:07 [PATCH 0/3] s390/pci: Updates to PCI insn tracing Gerd Bayer
2026-10-07 11:07 ` [PATCH 1/3] s390/pci: Adjust alignment in insn trace Gerd Bayer
2026-10-07 11:14   ` sashiko-bot
2026-10-07 13:10     ` Gerd Bayer
2026-10-07 14:36       ` Niklas Schnelle
2026-10-07 11:07 ` [PATCH 2/3] s390/pci: Use 32-bit fh outside of inline asm Gerd Bayer
2026-10-07 11:19   ` sashiko-bot
2026-10-07 14:46     ` Gerd Bayer
2026-10-07 16:21       ` Matthew Rosato
2026-10-08  8:50         ` Gerd Bayer [this message]
2026-10-08  9:28           ` Niklas Schnelle
2026-10-07 11:07 ` [PATCH 3/3] s390/pci: Add function handle to RPCIT insn trace Gerd Bayer
2026-10-07 11:17   ` sashiko-bot
2026-10-07 13:12     ` Gerd Bayer
2026-10-07 14:38       ` Niklas Schnelle

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=f082e43f024a3aa4a897b388b346eff4edea37ec.camel@linux.ibm.com \
    --to=gbayer@linux.ibm.com \
    --cc=agordeev@linux.ibm.com \
    --cc=borntraeger@linux.ibm.com \
    --cc=gor@linux.ibm.com \
    --cc=hca@linux.ibm.com \
    --cc=linux-s390@vger.kernel.org \
    --cc=mjrosato@linux.ibm.com \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=schnelle@linux.ibm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox