Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Anthony Krowiak <akrowiak@linux.ibm.com>
To: sashiko-reviews@lists.linux.dev
Cc: Christian Borntraeger <borntraeger@linux.ibm.com>,
	Alexander Gordeev <agordeev@linux.ibm.com>,
	Heiko Carstens <hca@linux.ibm.com>,
	Vasily Gorbik <gor@linux.ibm.com>,
	linux-s390@vger.kernel.org, kvm@vger.kernel.org,
	Matthew Rosato <mjrosato@linux.ibm.com>,
	Jason Herne <jjherne@linux.ibm.com>
Subject: Re: [PATCH v9 3/6] s390/vfio-ap: Fix unbounded loop in apq_reset_check()
Date: Tue, 29 Sep 2026 12:57:18 -0400	[thread overview]
Message-ID: <5c7cd376-2b86-45d9-b898-a8e927e2deb8@linux.ibm.com> (raw)
In-Reply-To: <20260929123815.287871F000FF@smtp.kernel.org>



On 9/29/26 8:38 AM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 6 potential issue(s) to consider:
> - [Critical] crypto: zeroization bypass and DMA vulnerability if ZAPQ returns BUSY
> - [Critical] crypto: Use-After-Free of vdev pointer after reset timeout
> - [High] crypto: resource leak and erroneous WARN when ZAPQ returns BUSY
> - [High] crypto: zeroization bypass if non-reset state change completes at timeout
> - [Medium] crypto: spurious ZAPQs from conflating TAPQ and ZAPQ BUSY responses
> - [Low] crypto: dead code setting fallback reset status on timeout
> --
>
> commit 625818127ace62cbd33b251880fe0f1f6d732f69
> Author: Anthony Krowiak <akrowiak@linux.ibm.com>
>
> s390/vfio-ap: Fix unbounded loop in apq_reset_check()
>
> This commit fixes an unbounded wait loop in the apq_reset_check() worker by
> adding a 2000ms timeout. It also updates TAPQ status checks to accurately
> verify asynchronous queue zeroization completion before freeing AQIC
> resources.
>
>> diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_ap_ops.c
>> index 087e8474a34ab..07fbfa6f10152 100644
>> --- a/drivers/s390/crypto/vfio_ap_ops.c
>> +++ b/drivers/s390/crypto/vfio_ap_ops.c
>> @@ -2164,6 +2164,12 @@ static int apq_status_check(int apqn, struct ap_queue_status *status)
>>   		return -EBUSY;
>>   
>>   	case AP_RESPONSE_BUSY:
>> +		/*
>> +		 * The queue is busy with something unrelated to a reset and our
>> +		 * ZAPQ was rejected outright. Re-issue the ZAPQ.
>> +		 */
>> +		return -EAGAIN;
>> +
> [Severity: Medium]
> Does returning `-EAGAIN` here trigger spurious ZAPQs by conflating TAPQ and
> ZAPQ BUSY responses?
>
> `apq_status_check()` evaluates the output of `ap_tapq()`. If `ap_tapq()`
> returns `AP_RESPONSE_BUSY`, it means the read instruction itself was
> transiently rejected. However, this code appears to incorrectly assume it
> means the previous ZAPQ was rejected outright, returning `-EAGAIN`.
>
> This forces the worker loop to immediately issue a new `ap_zapq()`, which
> overwrites `q->reset_status` and destroys the correct queue state tracking,
> likely leading to further failures.

Returning -EAGAIN causes apq_reset_check() to spuriously
issue a new ZAPQ and overwrite q->reset_status rather than continuing
to poll. We will change apq_status_check() to return -EBUSY on
AP_RESPONSE_BUSY.

>
> [ ... ]
>
>> @@ -2223,30 +2280,73 @@ static void apq_reset_check(struct work_struct *reset_work)
>>   			 */
>>   			memcpy(&q->reset_status, &status, sizeof(status));
>>   			return;
>> -		}
>> -		if (ret == -EBUSY) {
>> +		} else if (elapsed >= AP_RESET_MAX_WAIT) {
>> +			/*Timed out without being able to verify zapq completed */
>> +			if (!ret || ret == -ENODEV) {
> [Severity: High]
> Does evaluating the timeout condition before the retry logic allow an
> unrelated state change completion to bypass the required zeroization step?
>
> If the initial ZAPQ returns `AP_RESPONSE_STATE_CHANGE_IN_PROGRESS`, the worker
> polls until it completes before retrying ZAPQ. If the state change completes
> exactly at `elapsed >= AP_RESET_MAX_WAIT`, `ap_tapq()` indicates the queue is
> idle, and `ret = 0`.
>
> Because the loop evaluates the timeout condition `else if (elapsed >=
> AP_RESET_MAX_WAIT)` before the ZAPQ retry logic, the timeout block executes
> `apq_reset_finalize()`, forces the status to `AP_RESPONSE_NORMAL`, and returns.
>
> The required `ap_zapq()` is entirely bypassed, leaving the queue unzeroized
> while freeing AQIC resources, resulting in a DMA vulnerability.

Thinking this through it makes no sense to issue a PQAP-TAPQ for
AP_RESPONSE_STATE_CHANGE_IN_PROGRESS, because that response code is not 
valid for PQAP_TAPQ, so calling apq_status_check() will not tell us 
anything about the ZAPQ, because the response code indicates it failed 
to execute. The fix here is to check for 
AP_RESPONSE_STATE_CHANGE_IN_PROGRESS as well as AP_RESPONSE_BUSY right 
after the msleep and update of timeout counter and continue the loop.
>
>> +				/*
>> +				 * Zeroization confirmed (ret == 0): the TAPQ status bits
>> +				 * indicate the async portion of the ZAPQ completed
>> +				 * successfully. Free AQIC resources and return.
>> +				 *
>> +				 * Queue non-operational (ret == -ENODEV): the queue is
>> +				 * deconfigured or checkstopped; interrupts are not
>> +				 * possible so AQIC resources can be safely freed.
>> +				 * Zeroization cannot be confirmed in this state, but the
>> +				 * queue cannot generate interrupts, so the NIB page is
>> +				 * no longer a DMA target and it is safe to free it.
>> +				 */
>> +				apq_reset_finalize(q, &status, ret);
>> +				return;
>> +			}
>> +
>> +			report_apq_reset_check_timeout(q);
>> +
>> +			/*
>> +			 * Zeroization could not be confirmed; set
>> +			 * reset_status to AP_RESPONSE_RESET_IN_PROGRESS.
>> +			 * This is used internally to signal that the reset
>> +			 * did not complete, and ensures that if the queue
>> +			 * is reset again, the re-issue logic in
>> +			 * apq_reset_check() will re-issue the ZAPQ.
>> +			 */
>> +			q->reset_status.response_code = AP_RESPONSE_RESET_IN_PROGRESS;
> [Severity: Low]
> Is this assignment to `q->reset_status.response_code` dead code?
>
> The timeout logic assigns `AP_RESPONSE_RESET_IN_PROGRESS` to theoretically
> trigger a re-issue during the next reset. However, the next reset requires
> calling `vfio_ap_mdev_reset_queue()`, which immediately executes
> `status = ap_zapq(q->apqn, 0)` and unconditionally overwrites
> `q->reset_status` via `memcpy`.

Valid. Setting q->reset_status.response_code to
AP_RESPONSE_RESET_IN_PROGRESS on timeout is overwritten on any
subsequent reset because vfio_ap_mdev_reset_queue() immediately issues
ap_zapq() and overwrites q->reset_status with the new ZAPQ status. We
will remove this redundant assignment.

>
>> +
>> +			return;
> [Severity: Critical]
> Does returning early here skip cleanup and leave `q->aqic_resources.vdev`
> as a dangling pointer?
>
> `q->aqic_resources.vdev` points to the `vdev` embedded in `ap_matrix_mdev`
> without holding a reference count. When the new timeout mechanism returns
> early without calling `apq_reset_finalize()`, it leaves
> `q->aqic_resources.vdev` populated.
>
> When the mediated device is removed, `vfio_ap_mdev_remove()` flushes the
> worker and calls `vfio_put_device()`, which frees the `ap_matrix_mdev`.
> This turns `q->aqic_resources.vdev` into a dangling pointer on a surviving
> `vfio_ap_queue`.
>
> Later, if the AP queue is unbound or reused, `vfio_ap_free_aqic_resources()`
> unconditionally dereferences this pointer via
> `vfio_unpin_pages(q->aqic_resources.vdev, ...)`, which could trigger a
> use-after-free on the host.

Valid. When apq_reset_check() times out without freeing AQIC resources,
q->aqic_resources.vdev retains a raw pointer to matrix_mdev->vdev. If
the mdev is subsequently removed, matrix_mdev is freed while the queue
remains bound. A later call to vfio_ap_free_aqic_resources() dereferences
the stale vdev pointer in vfio_unpin_pages(), causing a use-after-free.
We will handle stale pointer tracking to avoid this dereference.

>
>> +		} else if (ret == -EBUSY) {
>>   			pr_notice_ratelimited(WAIT_MSG, elapsed,
>>   					      AP_QID_CARD(q->apqn),
>>   					      AP_QID_QUEUE(q->apqn),
>>   					      status.response_code,
>>   					      status.queue_empty,
>>   					      status.irq_enabled);
>> +			continue;
>>   		} else {
>> -			if (q->reset_status.response_code == AP_RESPONSE_RESET_IN_PROGRESS ||
>> -			    q->reset_status.response_code == AP_RESPONSE_BUSY ||
>> -			    q->reset_status.response_code == AP_RESPONSE_STATE_CHANGE_IN_PROGRESS ||
>> -			    ret == -EAGAIN) {
>> +			if (ret == -EAGAIN ||
>> +			    q->reset_status.response_code == AP_RESPONSE_RESET_IN_PROGRESS ||
>> +			    q->reset_status.response_code == AP_RESPONSE_STATE_CHANGE_IN_PROGRESS) {
> [Severity: Critical]
> Is it possible for `AP_RESPONSE_BUSY` to be missed in this retry check?
>
> If `ap_zapq()` inside the worker loop is rejected with `AP_RESPONSE_BUSY`,
> `q->reset_status.response_code` is set to `AP_RESPONSE_BUSY`.
> On the next iteration, if the queue becomes idle, `ap_tapq()` returns
> `AP_RESPONSE_NORMAL` (ret = 0).
>
> Since `AP_RESPONSE_BUSY` is missing from the retry condition, the worker seems
> to skip re-issuing the ZAPQ. It would fall through to `apq_reset_finalize()`,
> falsely overriding the status to `AP_RESPONSE_NORMAL` and freeing the AQIC
> resources. Can this leave secret keys intact and expose the host to wild
> DMA writes?

The AP_RESPONSE_BUSY was intentionally removed because it is not a valid 
response code from ZAPQ. If it is returned from ZAPQ, it will fall 
through to the default case in vfio_ap_mdev_reset_queue() because it is 
not architecturally possible to receive that response code for ZAPQ.
>
>> +				status = ap_zapq(q->apqn, 0);
>> +				memcpy(&q->reset_status, &status, sizeof(status));
>>   				continue;
>>   			}
>
> I also have a question about `vfio_ap_mdev_reset_queue()`:
>
> drivers/s390/crypto/vfio_ap_ops.c:vfio_ap_mdev_reset_queue() {
>      ...
> 	switch (status.response_code) {
> 	case AP_RESPONSE_NORMAL:
> 	case AP_RESPONSE_RESET_IN_PROGRESS:
> 	case AP_RESPONSE_STATE_CHANGE_IN_PROGRESS:
>      ...
> }
>
> [Severity: High]
> Does lacking `case AP_RESPONSE_BUSY:` here cause legitimately busy queues
> to fail in the default case?
>
> If the queue is transiently busy, ZAPQ returns `AP_RESPONSE_BUSY`. Execution
> falls to the `default:` branch, triggering a host WARN (falsely claiming a
> hardware bug) and intentionally leaking AQIC resources without queueing the
> worker to retry. Could this allow a guest that keeps the queue busy to cause
> resource leaks and kernel warnings on the host when the device is removed or
> reset?

See response to the previous question. Since AP_RESPONSE_BUSY is not a
valid response code for ZAPQ, it is in fact a hardware bug or an I/O
issue or something of the sort.

>


  reply	other threads:[~2026-09-29 16:57 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29 12:18 [PATCH v9 0/6] s390/vfio-ap: Fix pre-existing bugs in vfio_ap device driver Anthony Krowiak
2026-09-29 12:18 ` [PATCH v9 1/6] s390/vfio-ap: Fix leaks of pinned NIB and registered GISC Anthony Krowiak
2026-09-29 12:33   ` sashiko-bot
2026-09-29 13:38     ` Anthony Krowiak
2026-09-29 12:18 ` [PATCH v9 2/6] s390/vfio-ap: Fix failure to release IRQ notification eventfd contexts Anthony Krowiak
2026-09-29 12:25   ` sashiko-bot
2026-09-29 12:18 ` [PATCH v9 3/6] s390/vfio-ap: Fix unbounded loop in apq_reset_check() Anthony Krowiak
2026-09-29 12:38   ` sashiko-bot
2026-09-29 16:57     ` Anthony Krowiak [this message]
2026-09-29 12:18 ` [PATCH v9 4/6] s390/vfio-ap: Use AP_DOMAINS for adm_add bitmap size in vfio_ap_mdev_cfg_add() Anthony Krowiak
2026-09-29 12:27   ` sashiko-bot
2026-09-29 12:18 ` [PATCH v9 5/6] s390/vfio-ap: fix queue state leakage to guest and host Anthony Krowiak
2026-09-29 12:33   ` sashiko-bot
2026-09-29 12:18 ` [PATCH v9 6/6] s390/vfio-ap: replace guest-reachable WARNs with ratelimited warnings Anthony Krowiak
2026-09-29 12:41   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=5c7cd376-2b86-45d9-b898-a8e927e2deb8@linux.ibm.com \
    --to=akrowiak@linux.ibm.com \
    --cc=agordeev@linux.ibm.com \
    --cc=borntraeger@linux.ibm.com \
    --cc=gor@linux.ibm.com \
    --cc=hca@linux.ibm.com \
    --cc=jjherne@linux.ibm.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-s390@vger.kernel.org \
    --cc=mjrosato@linux.ibm.com \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox