Linux s390 Architecture development
 help / color / mirror / Atom feed
From: Matthew Rosato <mjrosato@linux.ibm.com>
To: Gerd Bayer <gbayer@linux.ibm.com>,
	sashiko-reviews@lists.linux.dev,
	Niklas Schnelle <schnelle@linux.ibm.com>
Cc: Alexander Gordeev <agordeev@linux.ibm.com>,
	Vasily Gorbik <gor@linux.ibm.com>,
	linux-s390@vger.kernel.org,
	Christian Borntraeger <borntraeger@linux.ibm.com>,
	Heiko Carstens <hca@linux.ibm.com>
Subject: Re: [PATCH 2/3] s390/pci: Use 32-bit fh outside of inline asm
Date: Wed, 7 Oct 2026 12:21:15 -0400	[thread overview]
Message-ID: <9e3894fa-0a1e-432b-ba1d-76ea34d4f648@linux.ibm.com> (raw)
In-Reply-To: <8086f631ce7ce60f1763171c8f05b050ff6c73fd.camel@linux.ibm.com>

On 10/7/26 10:46 AM, Gerd Bayer wrote:
> On Wed, 2026-10-07 at 11:19 +0000, sashiko-bot@kernel.org wrote:
>> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
>> - [High] s390/pci: infinite loop in zpci_refresh_trans on CC=2
>> --
>>
>> commit 1c8706bb963f0d551da7d8535f804ca9f9b4305a
>> Author: Gerd Bayer <gbayer@linux.ibm.com>
>>
>> s390/pci: Use 32-bit fh outside of inline asm
>>     
>> zpci_refresh_trans() declared the PCI function handle as a u64 type just
>> because the actual RPCIT machine instruction required that in a 64bit
>> register, while the function handle is architected as a 32bit type and
>> defined as such in the zpci_dev struct.
>>
>>> diff --git a/arch/s390/pci/pci_insn.c b/arch/s390/pci/pci_insn.c
>>> index 2ba787e2304b5..053afb3e7a990 100644
>>> --- a/arch/s390/pci/pci_insn.c
>>> +++ b/arch/s390/pci/pci_insn.c
>>> @@ -117,13 +117,14 @@ static inline u8 __rpcit(u64 fn, u64 addr, u64 range, u8 *status)
>>>  	return CC_TRANSFORM(cc);
>>>  }
>>>  
>>> -int zpci_refresh_trans(u64 fn, u64 addr, u64 range)
>>> +int zpci_refresh_trans(u32 fh, u64 addr, u64 range)
>>>  {
>>> +	u64 req = ZPCI_CREATE_REQ(fh, 0, 0);
>>>  	bool retried = false;
>>>  	u8 cc, status;
>>>  
>>>  	do {
>>> -		cc = __rpcit(fn, addr, range, &status);
>>> +		cc = __rpcit(req, addr, range, &status);
>>>  		if (cc == 2) {
>>>  			udelay(ZPCI_INSN_BUSY_DELAY);
>>
>> [Severity: High]
>> Does this retry loop risk causing an infinite loop or CPU stall when handling
>> large DMA spaces?
>>

This should really have been flagged as a pre-existing issue vs a new
regression, it is not a behavior introduced (or changed) by your patch.

>> When the s390x RPCIT instruction sets CC=2, it indicates that it cannot complete
>> the operation in the allotted time. It then updates the address and range
>> registers to reflect the remaining uncompleted portion.
> 
> I've re-read the relevant ISA documentation a couple times now. I don't
> find any mention of address and range being modified after RPCIT was
> executed - regardless of condition code or status.
> 
> Maybe @Niklas or @Matt have an idea where this allegation of "CC=2
> means partial completion and one should retry with an updated
> iova/range" comes from?

FWIW, it's not coming from QEMU -- we don't even have a CC=2 return case
for RPCIT today, and we never update the iova/range on return.

Perhaps, lacking access to the ISA documentation, it made an assumption
that the range would be updated to shorten the next RPCIT as this would
allow us to converge towards a smaller and smaller range each attempt to
potentially not hit the busy condition eventually - sounds logical anyway?

But I also find nothing that indicates the range will/should be updated
on CC=2 (or otherwise) from RPCIT; so it does just sound like a bad
assumption.

> 
>> However, in __rpcit(), addr_range is passed using an input-only constraint ("d"
>> instead of "+d"), and the arguments are passed by value:
>>
>> arch/s390/pci/pci_insn.c:__rpcit() {
>> 	union register_pair addr_range = {.even = addr, .odd = range};
>> 	...
>> 		: [addr_range] "d" (addr_range.pair)
>> }
>>
>> Because the hardware's partial progress updates to addr and range are discarded,
>> this loop continually re-invokes __rpcit() with the unmodified original
>> addr and range.

That matches your observation:  we indeed aren't taking updates to the
range in __rpcit().  But as you say: there aren't supposed to be any.

>>
>> If scanning the initial portion consistently exhausts the hardware time slice
>> (such as during IOMMU TLB flushes on devices with large DMA spaces in
>> zpci_refresh_all()), will the operation fail to advance and loop infinitely?
> 

That concern sounds valid at face value; if we tried something and it
couldn't be completed in time, from a linux perspective we are simply
trying it again without changing anything and hoping everything works
for the best this time.

> I see that all the PCI related instructions in this file may retry
> indefinitely (with delay) on CC=2. However, the architecture guarantees
> that at some point the instruction will end with any of the other
> condition codes.

I think that is the rub.  Sashiko couldn't possibly know that such an
architecture guarantee exists.

I suspect the question is also a bit theoretical: what happens if the
range is so big that it's impossible for the RPCIT to process it before
the CC2 trigger point due simply to how big the range provided is?  Then
you're guaranteed to hit it every time unless you 'chip away' at it by
shrinking the range after each attempt.

I would assume/hope that the SDMA-EDMA range limitation we impose on the
aperture size already ensures that it is easily possible to handle a
RPCIT from SDMA thru EDMA without tripping CC=2.
So then the CC=2 case becomes purely situational vs a guaranteed result
even for the largest possible RPCIT request.  That already makes it far
more reasonable to simply try it again.

So: assuming that architecture guarantee stands (we know RPCIT can
possibly complete for the largest possible range SDMA-EDMA, and we have
some architected guarantee that the loop will be broken otherwise e.g.
architecture says it won't keep giving us CC2s forever) then it sounds
like nothing to fix, but maybe worth a comment in a future patch that
explains these guarantees, based on what you find in the ISA.

I suppose we can consider whether linux should have its own redundancy
here or not in addition to those guarantees.  That sounds well beyond
the scope of this series and goes back to what I mentioned at the start:
 this is pre-existing bevahior and should not hold up this patch either way.

Thanks,
Matt

  reply	other threads:[~2026-10-07 16:21 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-07 11:07 [PATCH 0/3] s390/pci: Updates to PCI insn tracing Gerd Bayer
2026-10-07 11:07 ` [PATCH 1/3] s390/pci: Adjust alignment in insn trace Gerd Bayer
2026-10-07 11:14   ` sashiko-bot
2026-10-07 13:10     ` Gerd Bayer
2026-10-07 14:36       ` Niklas Schnelle
2026-10-07 11:07 ` [PATCH 2/3] s390/pci: Use 32-bit fh outside of inline asm Gerd Bayer
2026-10-07 11:19   ` sashiko-bot
2026-10-07 14:46     ` Gerd Bayer
2026-10-07 16:21       ` Matthew Rosato [this message]
2026-10-08  8:50         ` Gerd Bayer
2026-10-08  9:28           ` Niklas Schnelle
2026-10-07 11:07 ` [PATCH 3/3] s390/pci: Add function handle to RPCIT insn trace Gerd Bayer
2026-10-07 11:17   ` sashiko-bot
2026-10-07 13:12     ` Gerd Bayer
2026-10-07 14:38       ` Niklas Schnelle

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=9e3894fa-0a1e-432b-ba1d-76ea34d4f648@linux.ibm.com \
    --to=mjrosato@linux.ibm.com \
    --cc=agordeev@linux.ibm.com \
    --cc=borntraeger@linux.ibm.com \
    --cc=gbayer@linux.ibm.com \
    --cc=gor@linux.ibm.com \
    --cc=hca@linux.ibm.com \
    --cc=linux-s390@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=schnelle@linux.ibm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox